ISSN-registered · Peer-reviewed · Open Access
JournalsAboutContact
Scientific Writing

How to Present a Forest Plot and Heterogeneity in a Meta-Analysis

DE By Directive Editorial Team, Directive Publications ·4 Oct 2026 ·7 min read
How to Present a Forest Plot and Heterogeneity in a Meta-Analysis

A forest plot shows each study's effect estimate as a square sized by its weight, with a horizontal line for its 95% confidence interval (CI), and the pooled result as a diamond. To report one, name the effect measure, model and estimator, then give the between-study variance (tau²), I² and, with enough studies, a prediction interval.

How a meta-analysis differs from a systematic review is covered in our comparison of review types. This guide follows the PRISMA 2020 checklist for reporting and the Cochrane Handbook (version 6.5) for methods.

Label every part of the forest plot

Cochrane's example figure has these parts.

ElementWhat to show or state
Study rowsOrder by weight, year or risk of bias, and say which
Data columnsEvents/total per group, or mean, standard deviation and sample size per group
SquaresPoint estimates; area shows weight
Horizontal linesConfidence intervals, usually 95%
DiamondPooled estimate at the centre; tips mark its confidence interval
AxisLog scale for ratios, so 0.5 and 2 sit equally far from 1
"Favours" labelsWhich side favours which group; sides flip between good and harmful outcomes
WeightsPercentage shares, which depend on the model
Statistics lineHeterogeneity statistics and the test for overall effect

The vertical line marks no difference (1 for a ratio, 0 for a difference); our guide to confidence intervals and effect sizes explains what crossing it means. Add a table of study results, as PRISMA's explanation paper asks.

Choose fixed-effect or random-effects before you see I²

A fixed-effect, or common-effect, model assumes every study estimates the same effect. A random-effects model assumes true effects vary around an average; its confidence interval is wider than the fixed-effect one whenever I² is above zero, and smaller studies get relatively more weight.

Cochrane makes no universal recommendation, but says the choice "should never be made on the basis of a statistical test for heterogeneity". Switching to random effects when I² exceeds 50% is "strongly discouraged" in PRISMA's explanation paper. Fix the model in advance; this applies the rule in our research protocol guide that unplanned analyses are labelled exploratory. PRISMA 2020 item 13d and its explanation paper ask for:

  • the model, and why you chose it;
  • the method, such as Mantel–Haenszel or inverse variance;
  • for random effects, the tau² estimator, such as DerSimonian–Laird or restricted maximum likelihood (REML), and the interval method, such as Wald-type or Hartung–Knapp–Sidik–Jonkman (HKSJ);
  • how you assessed heterogeneity, such as visual inspection, Cochran's Q test, tau², I² and a prediction interval;
  • the software and version.

Cochrane notes that a large simulation favoured REML yet found no approach universally preferable. With few studies, Wald-type intervals can be too narrow and HKSJ intervals too wide. A random-effects model does not remove heterogeneity, and when smaller studies are biased it makes the bias worse.

Report tau², I² and a prediction interval together

  • Q test: give Cochran's Q (labelled Chi² in Cochrane's figure), its degrees of freedom (df) and P. Its power is low when studies are few or small, so a non-significant result does not show homogeneity.
  • I²: the percentage of variability due to heterogeneity rather than chance. It is very uncertain with few studies, so add its confidence interval.
  • Tau²: the between-study variance on the analysis scale, such as log risk ratios.
  • Prediction interval: where the true effect of a new, similar study is expected to lie. It is not a confidence interval for the average.

I² is a proportion, not a size: the same spread of true effects gives a higher I² when studies are larger and more precise. Tau, the square root of tau², shows that spread on the analysis scale, and the prediction interval shows it on the effect scale.

Higgins and colleagues (BMJ, 2003) only "tentatively" linked 25%, 50% and 75% to low, moderate and high. The Cochrane Handbook warns that thresholds "can be misleading" and offers overlapping bands for meta-analyses of randomised trials. Reading them depends on effect size, direction and the evidence for heterogeneity.

I²Cochrane's rough guide
0% to 40%might not be important
30% to 60%may represent moderate heterogeneity
50% to 90%may represent substantial heterogeneity
75% to 100%considerable heterogeneity

Cochrane encourages a prediction interval with about five or more studies and no clear funnel-plot asymmetry. Its 95% limits are M ± t × √(tau² + SE²), on the log scale for ratio measures. Here M is the pooled estimate, SE its standard error, and t the 97.5th percentile of the t-distribution with k−1 degrees of freedom, where k is the number of studies.

Look for causes before interpreting the average (PRISMA items 13e and 20c). Rule out extraction and unit-of-analysis errors, which can mimic heterogeneity. Compare subgroups with a formal test of the difference between them, not separate P values; keep meta-regression for ten or more studies, and run analyses with and without outlying studies rather than excluding them.

Test funnel-plot asymmetry only when 10 or more studies contribute

Why missing studies bias a pooled estimate is explained in our guide to reporting negative and null results; PRISMA items 14 and 21 cover assessing that risk. A funnel plot shows effect estimates against their standard errors and displays small-study effects. Cochrane lists five possible sources of asymmetry: non-reporting bias, inflated effects in smaller studies, true heterogeneity, artefact (odds ratios, for example, are correlated with their standard errors) and chance.

So asymmetry "should not be considered to be diagnostic" of non-reporting bias, and contour-enhanced plots may help tell the causes apart. As a rule of thumb, use an asymmetry test only with at least 10 studies of differing size. The original Egger test is not recommended for odds ratios or standardised mean differences; for odds ratios, Cochrane cites the Harbord and Peters tests. Name the test and give its exact P value.

Count every participant once and pool only similar studies

  • Shared control arm: entering "dose 1 versus placebo" and "dose 2 versus placebo" from one trial counts the placebo group twice and inflates precision. Cochrane recommends combining the intervention arms into one comparison.
  • Several reports of one study: collate them so the study is the unit (the authors' side is in our guide to disclosing overlapping patient cohorts).
  • Repeated time points and clusters: one time point per study per analysis; account for clustering.
  • Mixed designs: do not pool randomised with non-randomised studies, and for non-randomised studies prefer adjusted estimates.

Combine studies only when patients, treatments, comparators and outcomes match closely enough for one pooled figure to mean something clinically; otherwise, explain why you did not pool.

Map each element to its PRISMA 2020 item

PRISMA asks you to report methods and results; it mandates no forest plot or particular statistic. The EQUATOR Network also lists the checklist.

ItemWhat to report
12Effect measure for each outcome, such as risk ratio or mean difference
13cTable and graph methods; consider stating the basis for row order
13dModel, method, heterogeneity measures, software and rationale
13e, 13fExploring heterogeneity; sensitivity analyses
14, 15Assessing missing-results bias and certainty
19Per-study summary statistics and estimates with precision
20a, 20bCharacteristics and risk of bias of contributing studies; summary estimate, precision, heterogeneity, direction of effect
20c, 20dHeterogeneity investigations; sensitivity analyses
21, 22Missing-results bias; certainty, which unexplained heterogeneity lowers
Abstract 8Summary estimate, interval and favoured group

Copy the caption and Results templates

Replace each bracketed slot.

Figure [n]. [Outcome], [intervention] versus [comparator]: [k] studies, [N] participants. Squares show study [effect measures], sized by [model] weight; lines show 95% CIs; the diamond shows the pooled estimate and its 95% CI. [How any prediction interval is drawn.] The vertical line marks no effect; the axis is [log or linear]; values [left or right] favour [group]. Pooling: [method]; tau²: [estimator]; CI: [method]. Rows ordered by [criterion].

[k] studies ([N] participants): [events/total] with [intervention] and [events/total] with [comparator]. Pooled [effect measure] [estimate] (95% CI [lower] to [upper]; P = [value]; [model], [estimator], [CI method]), favouring [group]. Tau² = [value]; I² = [value]% (95% CI [lower] to [upper]); Q = [value], df = [value], P = [value]. 95% prediction interval [lower] to [upper]. [Funnel-plot assessment, or why none was done.]

The meta-analysis below is invented to illustrate the structure; its studies and numbers are not real data. Eight trials compare intervention X with usual care for infection within 30 days, a harmful outcome, so a risk ratio (RR) below 1 favours X.

StudyIntervention XUsual careRR (95% CI)Weight (random effects)
A12/15025/1480.47 (0.25 to 0.91)10.9%
B9/9510/970.92 (0.39 to 2.16)7.0%
C28/41058/4050.48 (0.31 to 0.73)18.9%
D10/607/621.48 (0.60 to 3.62)6.5%
E18/22034/2180.52 (0.31 to 0.90)14.2%
F27/30028/3020.97 (0.59 to 1.61)15.5%
G15/18026/1760.56 (0.31 to 1.03)12.2%
H22/25030/2550.75 (0.44 to 1.26)14.8%
Total141/1665218/16630.66 (0.49 to 0.90)100.0%

Weak: "X significantly reduced infection (p < 0.05), with low heterogeneity (I² = 29%)."

Strong: "Eight trials (3,328 participants): infection in 141/1,665 with intervention X and 218/1,663 with usual care. Pooled RR 0.66 (95% CI 0.49 to 0.90; P = 0.016; random effects, REML, HKSJ), favouring X. Tau² = 0.036; I² = 29% (95% CI 0% to 86%); Q = 10.27, df = 7, P = 0.17. 95% prediction interval 0.38 to 1.14. In a sensitivity analysis, the fixed-effect Mantel–Haenszel RR was 0.65 (95% CI 0.53 to 0.79). Funnel-plot asymmetry was not tested because fewer than 10 studies contributed."

Five checks. "Significantly" with "p < 0.05" hides the effect size, its confidence interval and the exact P value. I² of 29% falls in Cochrane's 0% to 40% band, but its 95% CI of 0% to 86% spans all four bands, so "low" misleads. The confidence interval excludes 1 but the prediction interval crosses it: in a new study like these, the true effect could be no benefit, or even harm. This I² comes from the REML tau², as Cochrane's Review Manager (RevMan) software reports it; (Q − df)/Q gives 32%, so say which you used. The similar fixed-effect estimate suggests small-study effects change little.

At Directive Publications, our article types page requires a completed PRISMA checklist and flow diagram for a systematic review and meta-analysis. Our author guidelines ask you to match the manuscript to the relevant EQUATOR checklist. Review is double-blind, with at least two independent expert reviewers sought, and our reviewer guidelines ask whether the statistics or analytical methods are appropriate and correctly applied. We set no separate rules for forest plots. To flag an error in this guide, report the problem to us.

Judging each included study's risk of bias before pooling, and carrying the judgements into the analysis, is covered in risk of bias assessment with RoB 2, ROBINS-I or Newcastle-Ottawa.

Frequently Asked Questions

What does the diamond at the bottom of a forest plot show?
The diamond shows the pooled estimate from the meta-analysis. Its centre is the point estimate, and its left and right tips are the limits of the confidence interval, usually 95%. The caption should name the model behind it, because fixed-effect and random-effects diamonds can differ when the studies disagree.
Is there an I² value above which studies are too different to pool?
No single I² value decides that. The Cochrane Handbook gives overlapping ranges rather than cut-offs, so 30% to 60%, for example, may represent moderate heterogeneity, and with few studies any I² value is very uncertain. Whether to pool is a judgement about whether the studies' patients, treatments, comparators and outcomes are alike enough to answer one clinical question, and about how far their results differ in size and direction.
Should I switch to a random-effects model if I² turns out to be high?
No. The PRISMA 2020 explanation and elaboration paper says switching models because I² exceeds 50% is strongly discouraged, and the Cochrane Handbook says the choice should never rest on a statistical test for heterogeneity. Choose the model in advance and give your reason. If you want to show what the other model changes, report it as a sensitivity analysis.
How is a prediction interval different from the confidence interval of a pooled estimate?
The confidence interval shows how precisely the average effect across studies has been estimated. The prediction interval estimates the range of true effects you could expect in another study like the ones included, so it widens as effects vary more between studies. A pooled estimate can have a confidence interval that excludes no effect while its prediction interval includes it.
DE
Directive Editorial Team
Directive Publications

The editorial team at Directive Publications — an international open-access publisher of peer-reviewed medical and scientific journals.

Publishing your research?

Directive Publications is an open-access publisher — every article peer-reviewed, Crossref-registered, and free to read under CC BY 4.0.

Submit a manuscript →Read our Scientific Writing policy →