A forest plot shows each study's effect estimate as a square sized by its weight, with a horizontal line for its 95% confidence interval (CI), and the pooled result as a diamond. To report one, name the effect measure, model and estimator, then give the between-study variance (tau²), I² and, with enough studies, a prediction interval.
How a meta-analysis differs from a systematic review is covered in our comparison of review types. This guide follows the PRISMA 2020 checklist for reporting and the Cochrane Handbook (version 6.5) for methods.
Label every part of the forest plot
Cochrane's example figure has these parts.
| Element | What to show or state |
|---|---|
| Study rows | Order by weight, year or risk of bias, and say which |
| Data columns | Events/total per group, or mean, standard deviation and sample size per group |
| Squares | Point estimates; area shows weight |
| Horizontal lines | Confidence intervals, usually 95% |
| Diamond | Pooled estimate at the centre; tips mark its confidence interval |
| Axis | Log scale for ratios, so 0.5 and 2 sit equally far from 1 |
| "Favours" labels | Which side favours which group; sides flip between good and harmful outcomes |
| Weights | Percentage shares, which depend on the model |
| Statistics line | Heterogeneity statistics and the test for overall effect |
The vertical line marks no difference (1 for a ratio, 0 for a difference); our guide to confidence intervals and effect sizes explains what crossing it means. Add a table of study results, as PRISMA's explanation paper asks.
Choose fixed-effect or random-effects before you see I²
A fixed-effect, or common-effect, model assumes every study estimates the same effect. A random-effects model assumes true effects vary around an average; its confidence interval is wider than the fixed-effect one whenever I² is above zero, and smaller studies get relatively more weight.
Cochrane makes no universal recommendation, but says the choice "should never be made on the basis of a statistical test for heterogeneity". Switching to random effects when I² exceeds 50% is "strongly discouraged" in PRISMA's explanation paper. Fix the model in advance; this applies the rule in our research protocol guide that unplanned analyses are labelled exploratory. PRISMA 2020 item 13d and its explanation paper ask for:
- the model, and why you chose it;
- the method, such as Mantel–Haenszel or inverse variance;
- for random effects, the tau² estimator, such as DerSimonian–Laird or restricted maximum likelihood (REML), and the interval method, such as Wald-type or Hartung–Knapp–Sidik–Jonkman (HKSJ);
- how you assessed heterogeneity, such as visual inspection, Cochran's Q test, tau², I² and a prediction interval;
- the software and version.
Cochrane notes that a large simulation favoured REML yet found no approach universally preferable. With few studies, Wald-type intervals can be too narrow and HKSJ intervals too wide. A random-effects model does not remove heterogeneity, and when smaller studies are biased it makes the bias worse.
Report tau², I² and a prediction interval together
- Q test: give Cochran's Q (labelled Chi² in Cochrane's figure), its degrees of freedom (df) and P. Its power is low when studies are few or small, so a non-significant result does not show homogeneity.
- I²: the percentage of variability due to heterogeneity rather than chance. It is very uncertain with few studies, so add its confidence interval.
- Tau²: the between-study variance on the analysis scale, such as log risk ratios.
- Prediction interval: where the true effect of a new, similar study is expected to lie. It is not a confidence interval for the average.
I² is a proportion, not a size: the same spread of true effects gives a higher I² when studies are larger and more precise. Tau, the square root of tau², shows that spread on the analysis scale, and the prediction interval shows it on the effect scale.
Higgins and colleagues (BMJ, 2003) only "tentatively" linked 25%, 50% and 75% to low, moderate and high. The Cochrane Handbook warns that thresholds "can be misleading" and offers overlapping bands for meta-analyses of randomised trials. Reading them depends on effect size, direction and the evidence for heterogeneity.
| I² | Cochrane's rough guide |
|---|---|
| 0% to 40% | might not be important |
| 30% to 60% | may represent moderate heterogeneity |
| 50% to 90% | may represent substantial heterogeneity |
| 75% to 100% | considerable heterogeneity |
Cochrane encourages a prediction interval with about five or more studies and no clear funnel-plot asymmetry. Its 95% limits are M ± t × √(tau² + SE²), on the log scale for ratio measures. Here M is the pooled estimate, SE its standard error, and t the 97.5th percentile of the t-distribution with k−1 degrees of freedom, where k is the number of studies.
Look for causes before interpreting the average (PRISMA items 13e and 20c). Rule out extraction and unit-of-analysis errors, which can mimic heterogeneity. Compare subgroups with a formal test of the difference between them, not separate P values; keep meta-regression for ten or more studies, and run analyses with and without outlying studies rather than excluding them.
Test funnel-plot asymmetry only when 10 or more studies contribute
Why missing studies bias a pooled estimate is explained in our guide to reporting negative and null results; PRISMA items 14 and 21 cover assessing that risk. A funnel plot shows effect estimates against their standard errors and displays small-study effects. Cochrane lists five possible sources of asymmetry: non-reporting bias, inflated effects in smaller studies, true heterogeneity, artefact (odds ratios, for example, are correlated with their standard errors) and chance.
So asymmetry "should not be considered to be diagnostic" of non-reporting bias, and contour-enhanced plots may help tell the causes apart. As a rule of thumb, use an asymmetry test only with at least 10 studies of differing size. The original Egger test is not recommended for odds ratios or standardised mean differences; for odds ratios, Cochrane cites the Harbord and Peters tests. Name the test and give its exact P value.
Count every participant once and pool only similar studies
- Shared control arm: entering "dose 1 versus placebo" and "dose 2 versus placebo" from one trial counts the placebo group twice and inflates precision. Cochrane recommends combining the intervention arms into one comparison.
- Several reports of one study: collate them so the study is the unit (the authors' side is in our guide to disclosing overlapping patient cohorts).
- Repeated time points and clusters: one time point per study per analysis; account for clustering.
- Mixed designs: do not pool randomised with non-randomised studies, and for non-randomised studies prefer adjusted estimates.
Combine studies only when patients, treatments, comparators and outcomes match closely enough for one pooled figure to mean something clinically; otherwise, explain why you did not pool.
Map each element to its PRISMA 2020 item
PRISMA asks you to report methods and results; it mandates no forest plot or particular statistic. The EQUATOR Network also lists the checklist.
| Item | What to report |
|---|---|
| 12 | Effect measure for each outcome, such as risk ratio or mean difference |
| 13c | Table and graph methods; consider stating the basis for row order |
| 13d | Model, method, heterogeneity measures, software and rationale |
| 13e, 13f | Exploring heterogeneity; sensitivity analyses |
| 14, 15 | Assessing missing-results bias and certainty |
| 19 | Per-study summary statistics and estimates with precision |
| 20a, 20b | Characteristics and risk of bias of contributing studies; summary estimate, precision, heterogeneity, direction of effect |
| 20c, 20d | Heterogeneity investigations; sensitivity analyses |
| 21, 22 | Missing-results bias; certainty, which unexplained heterogeneity lowers |
| Abstract 8 | Summary estimate, interval and favoured group |
Copy the caption and Results templates
Replace each bracketed slot.
Figure [n]. [Outcome], [intervention] versus [comparator]: [k] studies, [N] participants. Squares show study [effect measures], sized by [model] weight; lines show 95% CIs; the diamond shows the pooled estimate and its 95% CI. [How any prediction interval is drawn.] The vertical line marks no effect; the axis is [log or linear]; values [left or right] favour [group]. Pooling: [method]; tau²: [estimator]; CI: [method]. Rows ordered by [criterion].
[k] studies ([N] participants): [events/total] with [intervention] and [events/total] with [comparator]. Pooled [effect measure] [estimate] (95% CI [lower] to [upper]; P = [value]; [model], [estimator], [CI method]), favouring [group]. Tau² = [value]; I² = [value]% (95% CI [lower] to [upper]); Q = [value], df = [value], P = [value]. 95% prediction interval [lower] to [upper]. [Funnel-plot assessment, or why none was done.]
The meta-analysis below is invented to illustrate the structure; its studies and numbers are not real data. Eight trials compare intervention X with usual care for infection within 30 days, a harmful outcome, so a risk ratio (RR) below 1 favours X.
| Study | Intervention X | Usual care | RR (95% CI) | Weight (random effects) |
|---|---|---|---|---|
| A | 12/150 | 25/148 | 0.47 (0.25 to 0.91) | 10.9% |
| B | 9/95 | 10/97 | 0.92 (0.39 to 2.16) | 7.0% |
| C | 28/410 | 58/405 | 0.48 (0.31 to 0.73) | 18.9% |
| D | 10/60 | 7/62 | 1.48 (0.60 to 3.62) | 6.5% |
| E | 18/220 | 34/218 | 0.52 (0.31 to 0.90) | 14.2% |
| F | 27/300 | 28/302 | 0.97 (0.59 to 1.61) | 15.5% |
| G | 15/180 | 26/176 | 0.56 (0.31 to 1.03) | 12.2% |
| H | 22/250 | 30/255 | 0.75 (0.44 to 1.26) | 14.8% |
| Total | 141/1665 | 218/1663 | 0.66 (0.49 to 0.90) | 100.0% |
Weak: "X significantly reduced infection (p < 0.05), with low heterogeneity (I² = 29%)."
Strong: "Eight trials (3,328 participants): infection in 141/1,665 with intervention X and 218/1,663 with usual care. Pooled RR 0.66 (95% CI 0.49 to 0.90; P = 0.016; random effects, REML, HKSJ), favouring X. Tau² = 0.036; I² = 29% (95% CI 0% to 86%); Q = 10.27, df = 7, P = 0.17. 95% prediction interval 0.38 to 1.14. In a sensitivity analysis, the fixed-effect Mantel–Haenszel RR was 0.65 (95% CI 0.53 to 0.79). Funnel-plot asymmetry was not tested because fewer than 10 studies contributed."
Five checks. "Significantly" with "p < 0.05" hides the effect size, its confidence interval and the exact P value. I² of 29% falls in Cochrane's 0% to 40% band, but its 95% CI of 0% to 86% spans all four bands, so "low" misleads. The confidence interval excludes 1 but the prediction interval crosses it: in a new study like these, the true effect could be no benefit, or even harm. This I² comes from the REML tau², as Cochrane's Review Manager (RevMan) software reports it; (Q − df)/Q gives 32%, so say which you used. The similar fixed-effect estimate suggests small-study effects change little.
At Directive Publications, our article types page requires a completed PRISMA checklist and flow diagram for a systematic review and meta-analysis. Our author guidelines ask you to match the manuscript to the relevant EQUATOR checklist. Review is double-blind, with at least two independent expert reviewers sought, and our reviewer guidelines ask whether the statistics or analytical methods are appropriate and correctly applied. We set no separate rules for forest plots. To flag an error in this guide, report the problem to us.
Judging each included study's risk of bias before pooling, and carrying the judgements into the analysis, is covered in risk of bias assessment with RoB 2, ROBINS-I or Newcastle-Ottawa.