ISSN-registered · Peer-reviewed · Open Access
JournalsAboutContact
Scientific Writing

Statistical Reporting Done Right: p-values, Confidence Intervals and Effect Sizes

DE By Directive Editorial Team, Directive Publications ·6 Aug 2026 ·4 min read
Statistical Reporting Done Right: p-values, Confidence Intervals and Effect Sizes

Statistics are where many manuscripts are won or lost in review. Reviewers read the numbers closely, and nothing undermines confidence faster than a results section that leans on a single p-value and calls it a day. Reporting statistics well is not about advanced methods — it is about giving the reader enough, honestly, to judge what you found and to reproduce it. A few principles cover most of it.

Report the effect, not just the p-value

The most important number is usually the effect size — how big the difference or relationship actually is — because that is what matters in the real world. A study can be statistically significant and practically meaningless: with a large enough sample, a trivial difference clears p < 0.05. Always report the effect size (a mean difference, an odds ratio, a correlation) so readers can judge whether it matters, not just whether it reached a threshold.

Which effect size to report

OutcomeEffect sizeReport with it
Continuous (for example, blood pressure)Mean differenceGroup means with standard deviations, and the 95% confidence interval of the difference
Binary (for example, infection yes/no)Risk ratio or odds ratio, plus the risk differenceEvents and denominators in each group, and 95% confidence intervals
Time to eventHazard ratioEvents per group, length of follow-up, and the 95% confidence interval

For binary outcomes, give an absolute measure as well as a relative one. A risk ratio of 0.5 describes a fall from 2% to 1% and equally a fall from 40% to 20% — very different absolute changes.

Use the p-value for what it is

A p-value is widely misread. It is the probability of seeing data at least as extreme as yours if the null hypothesis were true. It is not the probability that your hypothesis is correct, and it says nothing about the size of an effect. Report exact values (p = 0.03, not "p < 0.05") except for very small ones, and never treat 0.049 and 0.051 as fundamentally different — they are not.

The American Statistical Association (ASA) made the same points in its 2016 statement on p-values. Its principles include that conclusions should not rest only on whether a p-value passes a threshold, and that proper inference requires full reporting of the analyses done.

  • Write "p < 0.001", never "p = 0.000". Software that prints 0.000 has rounded a small value.
  • Do not write "p = NS". Give the exact value.
  • State whether each test was one-sided or two-sided.

Give a confidence interval

A confidence interval shows the range of plausible values for your effect, and it does far more work than a p-value alone. "A 4.2-point reduction (95% CI 1.1 to 7.3)" tells the reader both the estimate and its uncertainty at a glance. Where you can, report the interval with every key estimate.

Word the interval carefully. A 95% confidence interval comes from a method that, across many repeated studies, would capture the true value 95% of the time. It does not mean there is a 95% probability that the true value lies inside this particular interval.

Use the interval to check your own tables. Take a two-sided p-value and a 95% interval from the same analysis. An interval that excludes the null value (0 for a difference, 1 for a ratio) goes with p < 0.05. A mismatch usually means a transcription error or two different tests.

Don't dichotomise the world

Splitting every result into "significant" and "not significant" throws away information and encourages bad inference. A non-significant result is not proof of no effect — it may just be an underpowered study. Describe your findings on a continuum: the size of the effect, the uncertainty around it, and what that means, rather than a binary verdict.

Two phrases draw reviewer objections. "A trend towards significance" describes a fixed p-value as if it were moving; report the estimate and its interval instead. Calculating "observed" power after the study adds nothing, because it follows directly from the p-value you already have. Our guide to reporting negative and null results shows how to write up such a finding.

Report enough to reproduce

Finally, give the mechanics. For every test, state which test you used, the sample size, the test statistic where relevant, and the exact p-value. If you ran many comparisons, say so and how you handled it. This detail is not padding — it is what lets a reader (or a reviewer) check your work, and it is central to reproducibility.

Describe the spread of roughly symmetrical data with the standard deviation (SD), and of skewed data with the median and interquartile range. The standard error (SE) describes the precision of an estimate, so use it only for estimates — or give a confidence interval instead. Label which one you used every time: "12.4 ± 3.1" with no label cannot be interpreted.

In a properly randomised trial, any difference between groups at baseline arose by chance. Significance tests on a baseline table therefore do not answer a useful question. Describe the groups instead. If a reviewer has already queried your analysis, see statistical reviewer comments, decoded.

Before and after: one result rewritten

The result below is invented to illustrate the structure; its numbers are not real data.

BeforeAfter
Pain was significantly lower in group A (p < 0.05).At 24 hours, mean pain score was 3.1 (SD 1.4) in group A (n = 40) and 4.0 (SD 1.6) in group B (n = 41). Mean difference -0.9 points (95% CI -1.6 to -0.2); p = 0.009, two-sided t-test.

The rewrite lets a reader judge whether a 0.9-point difference matters to patients, which the original did not.

A simple rule of thumb: for each key result, report the effect size, its confidence interval, and the exact p-value — in that order of importance.

For a time-to-event outcome, reporting Kaplan-Meier curves and hazard ratios covers censoring, median survival and the check of proportional hazards.

Updated 17 September 2026: We expanded this guide with an effect-size table, the ASA p-value principles, confidence interval checks, SD versus SE and a worked example.

Frequently Asked Questions

Is a p-value below 0.05 proof that my result is real?
No. A p-value is the probability of data at least this extreme if the null hypothesis were true — it is not the probability that your hypothesis is correct, nor a measure of effect size. Report it alongside an effect size and a confidence interval.
Should I report exact p-values or just "p < 0.05"?
Report the exact value (for example, p = 0.032) rather than a threshold, except for very small values where "p < 0.001" is conventional. Exact values let readers judge for themselves.
What is an effect size and why does it matter?
An effect size measures how big a difference or relationship is — the thing readers actually care about. A result can be statistically significant but trivially small, so always report the effect size, not just whether it was "significant".
DE
Directive Editorial Team
Directive Publications

The editorial team at Directive Publications — an international open-access publisher of peer-reviewed medical and scientific journals.

Publishing your research?

Directive Publications is an open-access publisher — every article peer-reviewed, Crossref-registered, and free to read under CC BY 4.0.

Submit a manuscript →Read our Scientific Writing policy →