ISSN-registered · Peer-reviewed · Open Access
JournalsAboutContact
Scientific Writing

Missing Data in Clinical Studies: Choosing and Justifying a Method

DE By Directive Editorial Team, Directive Publications ·4 Oct 2026 ·7 min read
Missing Data in Clinical Studies: Choosing and Justifying a Method

To choose a missing data method, judge why the values are missing: every method assumes a reason. Complete-case analysis and multiple imputation are defensible when their assumptions are plausible; last observation carried forward, mean imputation and missing indicators are not valid in general. Justify your assumption, then use a sensitivity analysis to test whether the conclusion survives a different one.

CONSORT 2025 item 21c asks how missing data were handled in a trial, and STROBE item 12(c) asks the same of observational studies. Counting missing values in each group, with reasons, is covered in how to write the Results section of a clinical paper.

Missing data mechanisms: MCAR, MAR and MNAR in plain words

The classification goes back to Rubin (1976). It describes how missing values relate to observed ones, not how many are missing.

MechanismWhat it meansInvented example
Missing completely at random (MCAR)No systematic difference between the missing values and the observed onesA freezer fault destroys a random batch of samples
Missing at random (MAR)Any systematic difference is explained by observed data; it does not mean "random"Older patients miss more visits, age is recorded, and within each age group the missing scores resemble the observed ones
Missing not at random (MNAR)Systematic differences remain after the observed data are taken into accountPatients whose pain worsens stop returning questionnaires, and nothing recorded predicts it

No test can show that data are MAR, because the values that would settle it are the ones you lack. Sterne and colleagues' 2009 paper calls missing at random "an assumption that justifies the analysis, not a property of the data". Record why each value is missing: the evidence is indirect. An administrative reason, such as a cancelled clinic, makes MAR more plausible; an unrecorded reason linked to worsening disease makes it less plausible.

When complete-case analysis is defensible, and what it costs

Complete-case analysis uses only the participants with every analysis variable recorded. It can be biased, and discarding incomplete records costs precision whenever they carry information about the question.

Complete-case analysis can be unbiased even when data are not MCAR (Hughes and colleagues, 2019). For most regression models, it is unbiased when the chance of being a complete case does not depend on the outcome once the model's covariates are taken into account. If only a once-measured outcome is missing and no auxiliary variables exist, imputation adds no information, and complete-case analysis is the better choice. For a small chart review, stating Hughes and colleagues' condition in the Methods, with the reason it is plausible, is what can justify the complete-case reply in ten statistical reviewer comments, decoded.

In a randomised trial, excluding any randomised participant means the analysis is no longer strictly intention-to-treat, as the CONSORT 2025 flow diagram guide explains.

Why last observation carried forward, mean imputation and missing indicators are not valid in general

Single imputation fills each gap with one value and analyses it as if measured, so standard errors are usually too small. Sterne and colleagues conclude that mean imputation, a missing-category indicator and carrying the last value forward are not statistically valid in general.

Last observation carried forward (LOCF) assumes the outcome stops changing when a participant leaves, which the CONSORT 2025 explanation and elaboration says "will rarely be valid". Nor is LOCF reliably conservative: depending on the disease course and the timing of dropout, it can bias a result either way, including in favour of a new treatment. A US National Research Council panel advised against single imputation as the primary approach unless its assumptions are scientifically justified.

A missing indicator, a "missing" category added to a variable, is valid for missing baseline covariates in a randomised trial but typically biased in non-randomised studies. Used as a confounder in an adjusted model, a "Not recorded" category becomes the missing-indicator method. Keeping "not documented" apart from "not present" is covered in writing the Methods for a retrospective chart review.

What a multiple imputation model must contain

Multiple imputation fills each gap with several plausible values drawn from a model of the observed data, analyses each completed dataset, and pools the results with Rubin's rules, whose standard error combines the uncertainty within each dataset with the variation between datasets. It is valid under MAR only if the imputation model is right, so check four things.

  • The outcome, even when only covariates are imputed. Leaving it out can weaken the associations you estimate.
  • Every analysis variable and interaction.
  • Auxiliary variables: variables outside the analysis model that predict the missing values or who is missing, such as an earlier outcome measurement. They can reduce bias and improve precision.
  • Suitable forms. Transform skewed variables or use predictive mean matching, and say how categorical variables were imputed.

The number of imputations is guidance, not a rule. Sterne and colleagues say at least 20 "may be preferable"; White, Royston and Wood (2011) proposed at least as many imputations as the percentage of incomplete cases. Under MNAR, multiple imputation can be as biased as a complete-case analysis, or more. Involve a statistician where you can, and report the complete-case result beside the imputed one.

Test the conclusion with a delta-adjusted or tipping-point analysis

Every analysis of incomplete data rests on assumptions the data cannot verify, so test the primary analysis with one that changes the assumption. A second method that also assumes MAR, such as a mixed model beside multiple imputation, does not count.

In a delta-adjusted analysis, you impute under MAR, then shift the imputed values by an amount, δ, so participants with missing data do worse, or better, than similar observed ones. δ = 0 is the MAR analysis. A tipping-point analysis increases δ until the conclusion changes, then asks whether that δ is clinically plausible. Cro and colleagues' practical guide cautions against tipping-point analyses when no careful thought has been given to which values of δ are plausible. Best-case and worst-case imputation suits a few missing binary outcomes, as in reporting a diagnostic accuracy study with STARD 2015.

Pre-specify the primary and sensitivity analyses, ideally in a registered protocol and certainly before unblinded data are seen. Agreement between them is reassuring, not proof. Box 8 of the CONSORT 2025 explanation and elaboration asks for at least a summary of the sensitivity analyses in the main paper, with full results in the supplement.

A decision table: situation, defensible method, what to write

Match your situation to a row; the last column is what your Methods should state. Repeated measures, clustering and complex imputation models still need a statistician.

SituationDefensible primary methodWhat to write
Values lost for a reason unrelated to any variableComplete-case analysisThe reason, and that the loss costs precision but is not expected to cause bias
Only a once-measured outcome missing; no auxiliary variablesComplete-case analysis, adjusting for predictors of missingnessThat being a complete case does not depend on the outcome given the covariates, and why that is plausible
Covariates missing, or auxiliary variables availableMultiple imputation under MARModel contents, number of imputations, complete-case comparison
Repeated outcome measurements with dropoutMultiple imputation or a likelihood-based mixed model; LOCF only if its assumption is justifiedThe MAR assumption and why it is plausible
Missing baseline covariates in a randomised trialMissing-indicator method or multiple imputationThat they are baseline covariates in a randomised comparison
Missing confounders in an observational studyMultiple imputation, or complete-case analysis if its condition holds; not a missing categoryWhich confounders had gaps, and why the assumption is plausible
Recorded reasons suggest MNAR, such as dropout after worsening symptomsThe most plausible stated assumption, plus delta-adjusted or tipping-point analysesThe δ range agreed in advance, and the tipping point
Treatment stopped, but the outcome can still be measuredNot missing data: keep measuringDiscontinuation and withdrawal reported separately, as the International Council for Harmonisation (ICH) E9(R1) addendum distinguishes them
The value cannot exist, such as quality of life after deathNot missing data: define its handling in advanceHow data truncated by death were handled

Template Methods sentences for a missing data paragraph

These sentences expand the single missing-data slot in the statistical analysis template of how to write the Methods section; keep the ones that match your method. The last one belongs in the Discussion.

Assumption. [variable] was missing for [n] of [N] participants ([%]), mainly because [reasons]. We assumed the data were missing at random given [observed variables], because [evidence].

Complete-case analysis. The primary analysis included the [n] participants with complete data. This is unbiased if being a complete case does not depend on [outcome] once [covariates] are accounted for, which we considered plausible because [reason].

Multiple imputation. [variables] were imputed by [method] in [software and version], using every analysis variable, the outcome, [interactions] and the auxiliary variables [list]. Estimates from [m] imputed datasets were pooled with Rubin's rules; complete-case results are reported for comparison.

Sensitivity analysis. Imputed values of [outcome] in [group] were shifted by δ from 0 to [value] [units] to find where [the conclusion] changed. Values up to [value] were judged plausible in [the statistical analysis plan, version and date].

Limitation (Discussion). Our estimates assume that missing [outcome] values were missing at random given [observed variables], which [recorded reasons or auxiliary variables] make [more or less] plausible. If participants with missing data in [group] had [worse or better] outcomes than similar observed participants, the [effect] would be [overestimated or underestimated].

At Directive Publications, our author guidelines ask you to match the manuscript to the relevant EQUATOR checklist. Review is double-blind: at least two independent expert reviewers are sought for each research manuscript, and the handling editor decides. Our reviewer guidelines list analysis integrity among the points a complete review usually considers: whether the statistics or analytical methods are appropriate and correctly applied. We have not published a separate statistical-review policy. If you find an error in this guide, here is how to report a problem to us.

Frequently Asked Questions

Can a statistical test show that my data are missing at random?
No. Whether data are missing at random or missing not at random depends on the values you did not observe, so no test on the observed data can settle it. You can make missing at random more plausible by recording why each value is missing and by including variables that predict missingness. A sensitivity analysis then shows how far the conclusion depends on the assumption.
Is last observation carried forward a conservative choice for a trial with dropouts?
Not reliably. Last observation carried forward assumes a participant's outcome stops changing after they leave, and depending on the course of the condition and the timing of dropout it can bias the result in either direction, including in favour of the new treatment. It also treats a carried value as a real measurement, so standard errors are usually too small. Use it as the primary analysis only if you can justify its assumption scientifically.
How many imputed datasets should a multiple imputation analysis use?
There is no fixed rule. Sterne and colleagues' 2009 guidance noted that five imputed datasets had been suggested as sufficient on theoretical grounds, but that at least 20 could be preferable to reduce sampling variability from the imputation process. White, Royston and Wood's rule of thumb links the number to the share of incomplete records: if 30% of records are incomplete, it suggests at least 30 imputations. State the number you used and why.
Is a complete-case analysis acceptable when only a small percentage of values is missing?
No percentage makes missing data safe to ignore. The European Medicines Agency's missing-data guideline for confirmatory trials says there is no rule on the maximum acceptable number of missing values, and methodologists advise against using the proportion missing to decide whether to use multiple imputation. A small proportion limits how far the result can move, but what matters is why the values are missing and whether that reason relates to the outcome, so state the reason and your assumption whatever the percentage.
DE
Directive Editorial Team
Directive Publications

The editorial team at Directive Publications — an international open-access publisher of peer-reviewed medical and scientific journals.

Publishing your research?

Directive Publications is an open-access publisher — every article peer-reviewed, Crossref-registered, and free to read under CC BY 4.0.

Submit a manuscript →Read our Scientific Writing policy →