What the service covers
Statistical analysis turns the data of a study into evidence about its question. The same data can be analyzed well or badly: a test can be chosen because it gives a small p-value and not because it fits the design, assumptions can go unchecked, missing data can be ignored, and the analysis can be repeated until it produces something publishable. A sound analysis is planned in advance, fits the design, checks its assumptions, reports what was done and shows how uncertain the results are.
This service supports researchers who have collected data and need an analysis that will withstand the scrutiny of reviewers and readers. It differs from the meta-analysis service, which combines the results of several studies, and from original research support, which also covers study design and manuscript preparation, although the services can be combined. The data themselves are yours and are not generated or altered by us, as set out in the research integrity statement.
What the service includes
- Review of the question and design to confirm what the data can and cannot answer.
- Statistical analysis plan that specifies the primary and secondary outcomes, the models, the handling of missing data and the planned subgroup or sensitivity analyses, written before the analysis.
- Data checking for entry errors, implausible values, duplicates and inconsistencies, with a log of every change.
- Descriptive analysis of the sample and the variables.
- Inferential analysis with models suited to the design and the outcome, and checks of their assumptions.
- Handling of missing data by an approach justified in the plan.
- Tables and figures formatted for journals.
- Methods and results text and analysis code that others can run.
- Response to statistical queries from reviewers, with re-analysis where requested.
The analysis plan
The statistical analysis plan is the most useful single safeguard in primary research. It states, before the analysis, which outcome is primary, what comparison answers the question, which model will be used and with which covariates, how missing data and outliers will be handled, which assumptions will be checked and what happens if they fail, and which analyses are secondary or exploratory. For randomized trials, guidance exists on what such a plan should contain. The plan does not prevent exploration; it separates the analyses that were planned from those that arose from looking at the data, so that the exploratory ones are reported as such.
When the data have already been seen, which is common, the plan is still written and the report is honest that it was written after data collection and, where it applies, after a first look. A plan made after seeing the results can still document the analytic reasoning, but it carries less weight, and the report says so.
Matching the analysis to the design
The design determines which analyses are valid. A few examples show what this means in practice.
- Randomized trials
- Analysis follows the randomization, usually by intention to treat. Baseline covariates used in stratification are included, and the effect is estimated with an appropriate model for the outcome: linear for continuous, logistic or log-binomial for binary, and Cox or parametric models for time to event. Clustered and crossover designs need analyses that account for them.
- Cohort and case-control studies
- Adjustment for confounding is the main task. Covariates are chosen from knowledge of the causal structure and not by significance, and the report describes the risk of residual confounding.
- Repeated measures and clustered data
- Observations from the same participant, family, clinic or school are correlated. Mixed-effects models or generalized estimating equations account for this, and ignoring it gives intervals that are too narrow.
- Survival data
- Time-to-event outcomes need methods that handle censoring. The proportional hazards assumption is checked and alternatives are used when it fails.
- Surveys and questionnaires
- Scale scores need reliability checks, and complex samples need weights. Ordinal items call for ordinal models and not for treatment as continuous by default.
- Diagnostic and prediction studies
- Accuracy is estimated with sensitivity, specificity and discrimination and calibration measures, and prediction models are validated and reported against TRIPOD.
Common models by type of outcome
The type of outcome narrows the choice of model. The table lists the usual starting points, which are adapted to the design, the sample size and the checks of assumptions.
| Outcome | Usual models | Effect measure |
|---|---|---|
| Continuous | Linear regression, analysis of covariance, linear mixed models | Mean difference, standardized mean difference |
| Binary | Logistic regression; log-binomial or Poisson with robust variance for risk ratios | Odds ratio, risk ratio, risk difference |
| Count or rate | Poisson or negative binomial regression | Rate ratio |
| Time to event | Cox proportional hazards, parametric survival models | Hazard ratio |
| Ordinal | Proportional odds and related models | Common odds ratio |
| Repeated or clustered | Mixed-effects models, generalized estimating equations | As for the outcome type |
Rank-based tests such as the Mann-Whitney or Kruskal-Wallis are options when assumptions of a parametric model cannot be met, but they test a different hypothesis and give no direct effect size, so a model with robust inference is often preferred. Whatever model is used, the report gives an estimate and an interval and not only a test statistic.
Sample size and power
The sample size determines whether a study can answer its question. A study with too few participants is likely to miss a real effect and, when it does find one, to overestimate it. Where the data are not yet collected, a sample-size calculation is part of design and requires an assumed effect size, variability, error rates and allowance for dropout, each of which should be justified from earlier evidence or from the smallest effect that would matter in practice. Where the data already exist, a post hoc calculation of power from the observed effect adds nothing, because it is a function of the p-value. It is more useful to report the confidence interval, which shows which effects the data rule out.
Missing data
Almost every dataset has missing values, and how they are handled can change the results. The first step is to understand why they are missing, whether by design, by chance, because of the value itself or because of other observed variables. A complete-case analysis, which drops participants with any missing value, is valid only under strong conditions and can lose power and introduce bias. Multiple imputation, which creates several plausible versions of the data and combines the results, is a widely used and defensible approach when the data are missing at random conditional on the variables in the imputation model. Where data may be missing not at random, sensitivity analyses explore how conclusions change under different assumptions. The approach is set in the plan, the amount of missing data is reported, and the conclusions are checked against alternatives.
Multiple testing and p-values
Each test carries a chance of a false positive, and many tests raise the chance that something will appear significant. The primary outcome is specified in advance for this reason, and secondary and exploratory analyses are labeled and, where appropriate, adjusted. A p-value is a statement about the compatibility of the data with a null hypothesis under a model, not the probability that a hypothesis is true or a measure of the size or importance of an effect. The statistical community's guidance is to report estimates with intervals, to avoid reducing results to a significant or non-significant verdict at a threshold, and to interpret findings in light of the design and the context. The analysis and the write-up follow this guidance.
Reproducibility and code
An analysis should be repeatable by someone else. The work is therefore done in scripts and not through manual steps, and the scripts, the software versions and the settings are recorded and delivered. Data cleaning is scripted, so that every change from the raw data is traceable. Random processes such as imputation and resampling use recorded seeds. Where data cannot be shared because of privacy, the code can still be provided, and a description of the data structure allows others to see what was done. See the guide on reproducible analysis.
Reporting
Reporting guidelines exist for most designs: CONSORT for randomized trials, STROBE for observational studies, STARD for diagnostic accuracy, TRIPOD for prediction models and others catalogued by the EQUATOR Network. They specify what a reader needs in order to judge a study. For statistical content, general guidance asks for a description of every method used, the software, how missing data were handled, the exact results with confidence intervals, and a clear separation of planned and exploratory analyses. The methods section is written to be understood by a statistician reviewer and by a subject-matter reader. See reporting guidelines.
Deliverables
- Statistical analysis plan and a log of any deviations from it.
- Data-checking report and a record of changes to the data.
- Results with tables and figures formatted for the target journal.
- Methods and results text describing the analysis in full.
- Analysis code and output logs with software versions.
- Support for reviewer comments on the statistics.
Get a quoteDescribe your design, outcome and target journal.
Optional extension: full manuscript and submission
Full manuscript and submission package
When the scope includes manuscript preparation and submission, the package also contains the following.
Full manuscript draft
Methods and results for your study built on the analysis, with the discussion drafted from your interpretation.
Reporting checklist
CONSORT, STROBE, STARD or the guideline that fits the design, completed against the final text.
Submission package
A journal recommendation, the manuscript formatted to that journal, a cover letter and supplementary files.
Revision round
Responses to statistical comments from reviewers, with re-analysis where requested.
Manuscript preparation follows the research integrity and authorship statement. The researchers who conceived the study and interpret its findings remain responsible for the content and its conclusions, and contributions that do not meet authorship criteria are acknowledged.
Limitations and boundaries
Statistics cannot repair a flawed design or poor data. A confounded comparison remains confounded after adjustment, a measurement that is invalid stays invalid, and a very small sample cannot support precise conclusions. We analyze the data we are given and do not collect, generate or alter data. If the data do not support the planned analysis, we say so and discuss the alternatives, which may include a different question. We will not select analyses to reach a preferred conclusion. Statistical analysis for a manuscript is research support and does not provide clinical advice, and the results are not patient-specific guidance.
This service provides research and evidence-synthesis support. It does not provide clinical advice.
How long does it take?
The time depends on the size and condition of the dataset, the number of outcomes and models, how much cleaning is needed, how complex the design is, and how many rounds of review follow. A clean dataset with a single primary outcome is a short project. A messy dataset with repeated measures and multiple outcomes takes longer, and the cleaning is often the larger part. After reviewing the data and the plan, a schedule and its dependencies are agreed.
Frequently asked questions
Do you collect data or run experiments?
No. The service analyzes data that you provide. It does not collect, generate or alter data.
Can you help before I collect data, with the sample size and design?
Design review and sample-size calculation are part of planning a study. They are most valuable before data collection, and the original research support service covers them as part of a wider project.
What if my results are not significant?
A non-significant result is still a result and is reported with its confidence interval. The analysis plan and the honest reporting of planned and exploratory analyses protect against the temptation to search for significance.
How do you handle missing data?
The approach is chosen from the pattern and likely cause of missingness and is set in the analysis plan. Multiple imputation is common when data are plausibly missing at random, with sensitivity analyses for other assumptions.
Do I get the code?
Yes. The analysis is done in scripts that are delivered with the software versions and settings, so that you or others can reproduce the results.
Is this clinical advice?
No. It is research support. The results of an analysis are not advice about the care of an individual patient.
References
- Gamble C, Krishan A, Stocken D, et al. Guidelines for the content of statistical analysis plans in clinical trials. JAMA. 2017;318(23):2337-2343.
- Schulz KF, Altman DG, Moher D; CONSORT Group. CONSORT 2010 statement: updated guidelines for reporting parallel group randomised trials. BMJ. 2010;340:c332.
- von Elm E, Altman DG, Egger M, Pocock SJ, Gotzsche PC, Vandenbroucke JP. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. BMJ. 2007;335:806-808.
- Bossuyt PM, Reitsma JB, Bruns DE, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015;351:h5527.
- Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. BMJ. 2015;350:g7594.
- Lang TA, Altman DG. Basic statistical reporting for articles published in biomedical journals: the SAMPL guidelines. Int J Nurs Stud. 2015;52(1):5-9.
- Wasserstein RL, Lazar NA. The ASA statement on p-values: context, process, and purpose. Am Stat. 2016;70(2):129-133.
- Wasserstein RL, Schirm AL, Lazar NA. Moving to a world beyond p < 0.05. Am Stat. 2019;73(sup1):1-19.
- Sterne JAC, White IR, Carlin JB, et al. Multiple imputation for missing data in epidemiological and clinical research: potential and pitfalls. BMJ. 2009;338:b2393.
- Little RJA, Rubin DB. Statistical Analysis with Missing Data. 3rd ed. Wiley; 2019.
- Harrell FE Jr. Regression Modeling Strategies. 2nd ed. Springer; 2015.
- Sandve GK, Nekrutenko A, Taylor J, Hovig E. Ten simple rules for reproducible computational research. PLoS Comput Biol. 2013;9(10):e1003285.