A subgroup analysis asks whether a treatment or exposure has a different effect in one kind of patient, such as older adults. To report one credibly, use a baseline variable and state a few hypotheses, with their direction, in advance. Test the interaction rather than comparing separate subgroup p-values, report every subgroup examined and label post hoc findings exploratory.
A real difference in effect between subgroups is called an interaction, or effect modification. A prognostic factor is different: it predicts the outcome whatever the treatment, and the Cochrane Handbook notes that confusing the two is common. Where the subgroup variable is sex, see our guide to the SAGER guidelines; this guide covers any baseline variable.
Define subgroups by baseline variables chosen in advance
Record the subgroup variable at baseline. The European Medicines Agency (EMA) guideline on subgroups in confirmatory trials warns that post-baseline variables may be affected by the treatment received. Comparing "responders" or "completers" is therefore no longer a randomised comparison. In a cohort, use variables measured at the start of follow-up.
Before the data are seen, record each hypothesis in the protocol or statistical analysis plan, with its rationale and expected direction. A 2012 review of 207 trials reporting subgroup analyses found 64 that claimed a subgroup effect on the primary outcome. Only 4 of those (6%) had correctly pre-specified its direction, and 54 (84%) met four or fewer of 10 credibility criteria.
An analysis plan entry might read:
We will examine [number] subgroup variables, each measured at baseline: [variables, with categories, or modelled as continuous]. We expect a [larger/smaller] effect in [subgroup] because [rationale or prior evidence]. Effect modification will be assessed with a test of interaction on the [ratio/difference] scale and reported as the difference between subgroup effects with its 95% confidence interval. Any other subgroup analysis will be reported as post hoc.
Test the interaction, not two separate p-values
The question is whether the effects in complementary subgroups, such as patients under 65 and those 65 or over, differ from each other. The CONSORT 2025 explanation for item 28 says a test of interaction is helpful here. If one is done, it asks for the difference between the subgroup effects with its confidence interval (CI), not just a p-value. One subgroup can be significant and its complement not, even with almost equal effects, so separate p-values cannot show a difference, as the SAGER guide illustrates. Overlapping CIs do not rule out an interaction either.
Expect the interaction test to have low power. With two equal subgroups, the difference between their effects has twice the standard error of the overall effect. In simulations by Brookes and colleagues, a trial with 80% power for its overall effect had only 29% power to detect an interaction of the same size. Detecting that interaction with 80% power needed about four times the sample size. A non-significant interaction is therefore weak evidence of a uniform effect.
State the scale on which effect modification was tested
An interaction depends on the scale. In this invented example, a 20% relative reduction takes risk from 30% to 24% in severe disease and from 10% to 8% in mild disease. There is no interaction on the relative scale, yet the absolute benefit is 6 percentage points against 2.
An interaction term in a logistic or Cox model, which estimates ratios, tests the multiplicative scale; comparing risk differences tests the additive scale. Name your scale and agree it with a statistician. The STROBE explanation and elaboration paper reports a consensus that the additive scale suits clinical decision making. Both kinds of effect are covered in reporting p-values, confidence intervals and effect sizes.
A quantitative interaction changes the size of an effect; a qualitative one reverses it, which the Cochrane Handbook calls rare. Yusuf and colleagues argued in 1991 that the overall result is usually a better guide to direction within a subgroup than the subgroup's own estimate.
Report every subgroup analysis you examined
CONSORT 2025 asks for subgroup methods in item 21d and results in item 28, each "distinguishing prespecified from post hoc". The item 28 explanation adds which subgroups were examined, why, and how many were pre-specified. STROBE items 12(b) and 17 cover observational studies, and SPIRIT 2025 item 27d covers trial protocols. Wording that labels an analysis pre-specified or exploratory is in our Methods section guide, and its place in the paper in our Results section guide.
Chance findings accumulate. Suppose no subgroup effect exists and k variables are each tested independently at p < 0.05. The chance of at least one false positive is then 1 − 0.95 to the power k.
| Subgroup variables tested (k) | Chance of at least one false positive | Expected false positives (0.05 × k) |
|---|---|---|
| 1 | 5.0% | 0.05 |
| 5 | 22.6% | 0.25 |
| 10 | 40.1% | 0.50 |
| 20 | 64.2% | 1.00 |
Real subgroup tests are not independent, so treat the percentages as a guide; the expected number holds regardless. The Cochrane Handbook, writing about meta-analyses, says it is difficult to suggest a maximum number of characteristics to examine. One safeguard is disclosure: report every subgroup, including those that showed nothing. Multiplicity corrections, and the cost of splitting a continuous outcome at a cut-off, are in our guide to choosing a statistical test.
A credibility checklist for a subgroup claim
The Instrument to assess the Credibility of Effect Modification Analyses (ICEMAN), from Schandelmaier and colleagues in 2020, is an appraisal tool, not a reporting guideline. It has five core questions for randomised trials and eight for meta-analyses, and the interaction p-value counts no more than any other item. Rows 1 to 5 quote its trial questions; row 6 adds a check from the sources named.
| Question | What a credible report shows, and where |
|---|---|
| 1. "Was the direction of effect modification correctly hypothesized a priori?" | The hypothesis and its direction, in a protocol or analysis plan (Methods) |
| 2. "Was the effect modification supported by prior evidence?" | Earlier studies or a biological rationale that predicted it (Introduction, Discussion) |
| 3. "Does a test for interaction suggest that chance is an unlikely explanation of the apparent effect modification?" | The interaction p-value and the difference between subgroup effects with its CI (Results) |
| 4. "Did the authors test only a small number of effect modifiers or consider the number in their statistical analysis?" | The number of subgroup variables examined, and any multiplicity allowance (Methods) |
| 5. "If the effect modifier is a continuous variable, were arbitrary cut points avoided?" | Cut-points fixed in advance, or the variable modelled continuously (Methods) |
| 6. Measured before randomisation (ICEMAN preliminary considerations; EMA) or, in a cohort, at the start of follow-up? | Baseline variables only (Methods) |
In a cohort, confounding that differs between subgroups can also produce an apparent difference, so give the adjusted estimate within each subgroup.
What a subgroup forest plot should show
A subgroup forest plot shows each subgroup's effect with its CI. Pocock and colleagues (2007) recommend:
- Squares of varying size, reflecting the amount of information.
- The overall estimate and its CI, with a dotted vertical line at the overall estimate.
- A log scale for ratio measures, labelled with which direction favours which treatment.
- Patients and events by arm, and one interaction p-value per variable rather than p-values within subgroups.
- No more detail than an exploratory subgroup analysis warrants.
The EMA guideline adds that a subgroup's complement should be shown with it, and that a forest plot does not determine credibility.
Word subgroup findings as exploratory unless they pass the checks
Keep a subgroup result out of the Abstract unless it was pre-specified and backed by an interaction test. The explanation for CONSORT 2025 item 1b warns against reporting only the statistically significant subgroup analyses. The EMA guideline adds that a subgroup used to rescue a formally failed trial allows no further confirmatory conclusions. The rule that an abstract claims no more than the main text is in matching the conclusion to your data.
These examples are invented; their numbers are not real data.
| Situation | Before: overclaims | After: labelled and supported |
|---|---|---|
| Pre-specified subgroup, separate p-values | Treatment reduced readmission at age 65 or over (risk ratio 0.70, 95% CI 0.52 to 0.94; p = 0.02) but not in younger patients (0.85, 95% CI 0.60 to 1.20; p = 0.36). | In a pre-specified analysis, the effect did not differ clearly by age: ratio of risk ratios 0.82 (95% CI 0.52 to 1.30; interaction p = 0.40). The wide interval does not exclude a substantial difference. |
| Post hoc finding among many | Benefit was confined to patients with diabetes (p = 0.04). | Diabetes was one of 12 subgroup variables examined post hoc. The effect appeared larger with diabetes than without: ratio of risk ratios 0.62 (95% CI 0.39 to 0.98; interaction p = 0.04). With 12 independent tests and no true differences, the chance of at least one p-value below 0.05 is about 46%, so the finding is exploratory. |
| Cohort, multiplicative scale | The association was stronger at age 75 or over (odds ratio 2.1) than below 75 (odds ratio 1.3). | The adjusted odds ratio was 2.1 (95% CI 1.2 to 3.8) at 75 or over and 1.3 (95% CI 0.8 to 2.1) below 75. On the multiplicative scale, the ratio of odds ratios was 1.62 (95% CI 0.76 to 3.43; interaction p = 0.21), so the data fit both no difference by age and a large one. |
Answer a reviewer who asks for more subgroups
You may run a subgroup a reviewer requests, label it post hoc and add it to your count, or decline in writing if there is no rationale or the groups are too small. If a reviewer challenges your claim, work through the checklist and lower the wording to match. In a small retrospective study, outcome counts within each subgroup, with no claim of a difference, are the honest substitute; see our guide to statistical reviewer comments.
Directive Publications has no published policy on subgroup analyses. Our author guidelines ask you to match the manuscript to the relevant EQUATOR Network checklist, such as CONSORT or STROBE. Our reviewer guidelines ask whether claims are justified by the data presented, without overstatement. Review is double-blind, at least two independent expert reviewers are sought for each research manuscript, and the handling editor decides. If you find an error in this guide, please report the problem to us.