A sample size calculation is done before a study begins: you name the primary outcome, state the smallest difference that would change clinical practice, estimate the variability you expect, fix the significance level and the power you will accept, then derive the number of participants. Calculating power afterwards from your own observed result is uninformative. Report confidence intervals instead.
What a sample size calculation is, and the six inputs it needs
A sample size calculation is an arithmetic statement of how many participants a defined statistical test needs to detect a difference of a defined size, at a defined error rate, if that difference is real. It is a planning tool that belongs in the protocol.
| Input | What it means | Where the number comes from |
|---|---|---|
| Primary outcome | The single measurement the study is designed to answer | Your research question. One outcome only |
| Minimum clinically important difference | The smallest change that would alter how a patient is treated | Clinical judgement or a published threshold |
| Expected variability | Standard deviation for a continuous outcome, or the control-group event rate for a binary one | A comparable published population, or your own audit data |
| Significance level (alpha) | The false positive rate you accept, conventionally 0.05, two-sided | Convention, stated explicitly |
| Power (1 minus beta) | The chance of detecting that difference if it truly exists, conventionally 80 or 90 per cent | Convention, stated explicitly |
| Attrition allowance | Inflation for loss to follow-up or unusable records | Your own drop-out experience |
One outcome drives the number, not all of them
A study powered for length of stay is not powered for mortality. Label every other outcome as secondary in the Methods: it pre-empts an obvious objection.
Why post-hoc power is uninformative, not merely discouraged
Post-hoc power, also called observed or retrospective power, is power recalculated after a study has finished, using the effect size that study observed. It answers no question you actually have.
The reason is arithmetic, not stylistic. For a given test and sample size, observed power is a direct function of the p-value from that same test: a large p-value always produces a low observed power, and a p-value just above the threshold always produces an observed power near or below 50 per cent. The two numbers carry identical information: observed power beside a p-value simply repeats it. It cannot tell you whether a non-significant result means the effect is absent or the study was too small, because it is computed from the result being questioned.
That is why "we performed a post-hoc power analysis, which showed the study was underpowered" is circular: our p-value was large, therefore our p-value was large.
When a reviewer writes that a study "appears underpowered", the real question is which effect sizes the data can still rule out. That question has an answer, and it is not a power figure.
Report the confidence interval instead of retrospective power
A confidence interval states the range of true effects compatible with your data. In a small clinical study it is the honest instrument, because it makes imprecision visible.
Worked example. A single-centre study compares a modified technique with standard care in 48 patients: 3 complications of 25 (12 per cent) against 5 of 23 (22 per cent). That is about 10 percentage points in favour of the new technique, with a 95 per cent confidence interval from roughly 31 points of benefit to 11 points of harm. Two sentences follow from it, and only one is true.
- Untrue: "There was no difference between the groups (p = 0.36)."
- True: "The complication rate was lower in the intervention group, but the 95 per cent confidence interval spans both a substantial reduction and a clinically important increase, so this study cannot distinguish benefit from harm."
A narrow interval — 2 points of benefit to 3 of harm — would instead support a claim that any remaining difference is small. Interval width, not the presence of an asterisk, is what a small study should be judged on. Our post on p-values, confidence intervals and effect sizes covers the conventions.
A sample size reporting template for your Methods
Paste the sample size template below into the Methods and replace every bracketed field. Each field exists because a reader cannot check the number without it.
Sample size. The primary outcome was
[outcome and how it was measured]. Based on[cited study, local audit, or stated assumption], we assumed[control-group event rate, or mean and standard deviation]. We considered[minimum clinically important difference, with units]the smallest difference worth detecting, because[what would change if it were real]. Using[named test, and the software or table used], with a two-sided alpha of[0.05]and power of[80 or 90] per cent,[n]participants per group were required. Allowing[x] per centfor[attrition], the recruitment target was[N].
Filled in, with a hypothetical surgical example:
Sample size. The primary outcome was surgical site infection within 30 days, recorded using the standard wound assessment in our unit. Based on our departmental audit for 2023 to 2024, we assumed a control-group infection rate of 20 per cent. We considered an absolute reduction of 10 percentage points the smallest difference worth detecting, because a smaller reduction would not change our prophylaxis protocol. Using a two-group chi-squared test with a two-sided alpha of 0.05 and 80 per cent power, 199 participants per group were required. Allowing 10 per cent for incomplete records, the recruitment target was 444, or 222 per group.
Note what the filled version admits: the assumption came from local audit data, not the literature.
Honest wording when the study is smaller than planned
Under-recruitment is a fact about a study, not a flaw in its authors. It belongs in the Results and the limitations, not in a retrospective calculation. Three wordings follow, one for each of three situations.
You calculated a target and did not reach it.
Recruitment was planned for
[N]participants but closed at[n]after[unit closure, slower eligibility than expected, funding end date]. The study therefore has lower precision than planned. Effect estimates are reported with 95 per cent confidence intervals, and no conclusion of equivalence is drawn from a non-significant result.
You did not calculate a target, because the sample was whatever existed.
No prospective sample size calculation was performed. All
[n]consecutive patients meeting the eligibility criteria between[dates]were included, so the sample size was determined by the size of that cohort. Estimates are reported with 95 per cent confidence intervals, and the study should be read as[hypothesis-generating or descriptive]rather than as a test of efficacy.
The study is a pilot or feasibility study.
This was a
[pilot or feasibility]study designed to estimate[recruitment rate, protocol adherence, outcome variability]ahead of a definitive study. It was not powered to test effectiveness, and no hypothesis test of the clinical outcome is reported. The observed[standard deviation or event rate]is presented to inform a future sample size calculation.
The third wording does something the other two cannot. A small study that reports the variability it observed hands the next investigator a real input, as our post on negative and null results sets out.
Where reporting guidelines ask about sample size
The EQUATOR Network hosts the checklists below. Match the manuscript to the one that fits your design before submitting.
| Design | Checklist | What it asks about sample size |
|---|---|---|
| Randomised trial | CONSORT 2025 item 16a | How sample size was determined, including all assumptions supporting the calculation |
| Trial protocol | SPIRIT 2025 | The estimated number of participants, how it was determined and the assumptions behind it, stated in advance |
| Pilot or feasibility trial | CONSORT extension for pilot and feasibility trials | The rationale for the number chosen, given no power claim |
| Cohort, case-control, cross-sectional | STROBE item 10 | How the study size was arrived at, including when data availability fixed it |
| Case report or case series | CARE for a case report, PROCESS for a surgical case series | Neither has a sample size item. Do not construct one |
The ICMJE Recommendations ask that statistical methods be described in enough detail for a knowledgeable reader to judge them. A sample size paragraph naming its assumptions meets that standard. A retrospective power figure gives a reader nothing to check.
Three checks to run before you submit
- Alpha, power, the assumed rate or standard deviation, its source and the test are all stated.
- The achieved sample is compared with the planned sample, with the reason for any shortfall.
- Every estimate carries a confidence interval, no observed power figure appears anywhere, and no sentence claims "no difference" on a non-significant p-value alone.
Sample size honesty is a writing problem more than a statistical one: it means committing to assumptions you would rather leave vague. Our Author Guidelines ask you to match your manuscript to the relevant EQUATOR checklist. See also how to write a limitations section, writing the Methods for a retrospective chart review and types of research articles explained.