ISSN-registered · Peer-reviewed · Open Access
JournalsAboutContact
Scientific Writing

How to Build Table 1: Presenting Baseline Characteristics

DE By Directive Editorial Team, Directive Publications ·27 Sep 2026 ·7 min read
How to Build Table 1: Presenting Baseline Characteristics

Table 1 baseline characteristics describe who was studied, with one column per comparison group and an optional total column. Choose the rows before analysis: demographics, the prognostic factors and confounders named in your Methods, and whatever readers need to judge applicability. Summarise each variable descriptively, show every category and its missing values, and omit p-values in a randomised trial.

Table 1 is where a reader checks whether your participants resemble their own patients and whether the groups started from a similar place. For a randomised trial, CONSORT 2025 item 25 asks for "a table showing baseline demographic and clinical characteristics for each group". For an observational study, STROBE item 14(a) asks for participant characteristics "and information on exposures and potential confounders", by group where applicable (STROBE on the EQUATOR Network).

How the Results text points to the table is covered in how to write the Results section; this guide builds the table itself.

Choose Table 1 baseline characteristics from your Methods, not your data

Every row needs a reason you could have stated before seeing the results. Draw the list from the protocol or analysis plan.

  • Who the participants are. Age, sex and the demographic variables your setting makes relevant.
  • How ill they were. Severity, stage and comorbidity, defined as in the Methods.
  • What predicts the outcome. Known prognostic factors, any variable used to stratify randomisation, and the outcome's baseline value if you measured it.
  • What you adjusted for. In an observational study, every confounder in the adjusted model.

Do not add or drop rows once you have seen which ones differ. The STROBE explanation and elaboration paper says confounders should be chosen from knowledge of, or explicit assumptions about, causal relations, not from p-values or statistical significance alone. When the list grows long, keep the variables that matter for interpreting the main result and move the rest to a supplementary table.

Summarise each baseline variable descriptively, never inferentially

Describe roughly symmetrical continuous variables by the mean and standard deviation (SD), and skewed ones by the median with the 25th and 75th percentiles, the limits of the interquartile range (IQR). That choice, and why a trial's baseline table carries no significance tests, is explained in our guide to p-values, confidence intervals and effect sizes. Three formatting points are specific to Table 1.

Show every level of a categorical variable with three or more levels, so its rows add up to the column total; a binary variable needs one row, such as "Women, n (%)". Report an ordered variable with a few categories, such as a disease stage, as numbers and proportions for each category, as the CONSORT 2025 explanation paper advises.

Give each group a column and put its size in the header

Use one column per comparison group, in the order the Methods introduce them, with each group's size in the header: "Day case (n = 84)". The header numbers must match the flow diagram. In a case-control study, the columns are cases and controls. A total column is optional; it helps readers who want one description of the whole sample.

Name the statistic and unit in the row label, not in every cell: "Age, years, mean (SD)". When the header gives each group's size, n (%) is enough; switch to n/N (%) wherever missing values change the denominator. General rules on counts, denominators and alignment are in preparing figures and tables for publication.

Record missing baseline values in the row where they occur

STROBE item 14(b) asks for the number of participants with missing data for each variable of interest. Because item 14 is starred, that count is given for each group where applicable. Put it where the reader needs it: a "Not recorded" row under the variable, or n/N in the row itself. A footnote saying that some data were missing does not meet the item. The reply line a reviewer expects once Table 1 shows missingness per variable is quoted in ten statistical reviewer comments, decoded.

Keep "Not recorded" separate from "No". A diabetes row that counts undocumented patients as non-diabetic can understate the prevalence, and it hides the gap.

Judge a baseline imbalance by its size, not by a p-value

In a randomised trial, proper random assignment prevents selection bias but does not guarantee similar groups, and the CONSORT 2025 explanation paper says the chance differences that remain should not be tested for significance. Judge each by how strongly the variable predicts the outcome and how large the imbalance is. A small gap in a strong prognostic factor can matter more than a large gap in a weak one, and deserves a sentence in the Discussion.

Never write that randomisation "failed" because a baseline difference crossed p < 0.05. In an observational study the groups are expected to differ, and the STROBE explanation paper also advises against significance tests in descriptive tables.

Show balance in an observational study with standardised mean differences

A standardised mean difference (SMD) expresses the gap between two groups in units of their pooled standard deviation. For a continuous variable, divide the difference in means by the square root of the average of the two variances. For a binary variable, divide the difference in proportions by the square root of the average of p(1 − p) in the two groups. Unlike a p-value, an SMD does not depend on the sample size, and because it has no units, an imbalance in age can be set beside an imbalance in diabetes.

In the invented example below, mean age was 52.3 years (SD 13.1) for day cases and 58.9 years (SD 14.2) for overnight stays. The difference is 6.6 years, the pooled SD is √((13.1² + 14.2²)/2) = 13.7, and the SMD is 6.6/13.7 = 0.48. For diabetes, 8.3% against 18.3% gives an SMD of 0.30.

Methodological papers on propensity-score matching use the SMD, not significance tests, to check balance in matched samples (an introduction to propensity-score methods). An SMD below 0.1 "has been taken to indicate a negligible difference", but the same paper notes that no threshold for important imbalance is universally agreed. No SMD can show balance in a confounder that was never measured. State in the footnote how any SMD was calculated.

A Table 1 template to copy, and a filled example

Copy the template and replace each slot. The SMD column is optional; delete it if you will not use it.

Characteristic[group a] (n = [n])[group b] (n = [n])Total (n = [n])SMD
[variable], [unit], mean (SD)[mean] ([sd])[mean] ([sd])[mean] ([sd])[smd]
[skewed variable], [unit], median (IQR)[median] ([q1] to [q3])[median] ([q1] to [q3])[median] ([q1] to [q3])[smd]
Not recorded, n[n][n][n]
[binary characteristic], n (%)[n] ([%])[n] ([%])[n] ([%])[smd]
[categorical variable], n (%)
[level 1][n] ([%])[n] ([%])[n] ([%])[smd]
[level 2][n] ([%])[n] ([%])[n] ([%])[smd]
Not recorded[n][n][n]
Table [x]. Baseline characteristics of [population], by [grouping variable]. Values are mean (SD), median (IQR, 25th to 75th percentile) or n (%). Percentages are of participants with a recorded value; rows marked "Not recorded" give the number with missing data. SMD, absolute standardised mean difference, calculated as [method]. [abbreviation], [full term].

The cohort below is invented to illustrate the structure; its numbers are not real data. It describes 210 adults having elective hernia repair at one hospital, 84 as day cases and 126 with an overnight stay.

CharacteristicDay case (n = 84)Overnight stay (n = 126)Total (n = 210)
Age, years, mean (SD)52.3 (13.1)58.9 (14.2)56.3 (14.1)
Women, n (%)31 (36.9)38 (30.2)69 (32.9)
Body mass index, kg/m², median (IQR)*27.1 (24.3 to 30.2)28.4 (25.0 to 32.6)27.8 (24.6 to 31.5)
Not recorded, n3912
Diabetes, n (%)7 (8.3)23 (18.3)30 (14.3)
Current smoker, n (%)15 (17.9)22 (17.5)37 (17.6)
American Society of Anesthesiologists (ASA) physical status, n (%)
ASA I38 (45.2)30 (23.8)68 (32.4)
ASA II41 (48.8)71 (56.3)112 (53.3)
ASA III5 (6.0)25 (19.8)30 (14.3)

Values are mean (SD), median (IQR) or n (%). *Median of patients with a recorded value: 81 day cases and 117 overnight stays. Percentages may not total 100 because of rounding.

Check Table 1 against the rest of the paper before you submit

  • Each header n matches the flow diagram.
  • Every row variable is defined in the Methods, with its unit and timing.
  • Each categorical variable's rows, including "Not recorded", add up to the column n.
  • Any total column equals the sum of the groups.
  • No p-values in a trial's table; no standard errors or confidence intervals for spread anywhere.
  • Decimal places are consistent within each row.
  • Every abbreviation is defined in the footnote.
  • The Results text names only the imbalances that matter, not every row.

Editors can also compare baseline tables across papers. Two papers from one centre with matching age, sex and comorbidity distributions suggest a shared cohort, and the overlap must be disclosed, as set out in disclosing overlapping patient cohorts. If you find an error in a table in an article we have published, here is how to report a problem to us.

Frequently Asked Questions

Which variables belong in Table 1 of an observational study?
STROBE item 14(a) asks for characteristics of participants, for example demographic, clinical and social ones, and "information on exposures and potential confounders". Because item 14 is starred, give them separately for cases and controls in a case-control study and, where applicable, for exposed and unexposed groups in a cohort or cross-sectional study. Include every confounder in your adjusted model, chosen from causal knowledge rather than from p-values.
Should Table 1 include a total column for all participants?
A total column is optional. It is useful when readers want a single description of the whole sample to compare with their own patients, and its n should equal the sum of the group columns. In a trial with several arms it can crowd the table, and the group columns remain the ones that carry the comparison.
Is a standardised mean difference below 0.1 proof that two groups are balanced?
No. Values below 0.1 are conventionally read as a negligible difference, but no threshold for important imbalance is universally agreed. A standardised mean difference also describes only the variables in the table, so it cannot rule out confounding by anything that was not measured.
Should a baseline table show mean ± SD or mean (SD)?
Write mean (SD), for example 52.3 (13.1), as the SAMPL (Statistical Analyses and Methods in the Published Literature) guidelines recommend. Name the statistic in the row label, so readers never have to guess whether a number is a standard deviation, a standard error or a range. Keep standard errors and confidence intervals out of baseline rows: they are inferential statistics about how precisely something is estimated, not descriptions of how participants vary.
DE
Directive Editorial Team
Directive Publications

The editorial team at Directive Publications — an international open-access publisher of peer-reviewed medical and scientific journals.

Publishing your research?

Directive Publications is an open-access publisher — every article peer-reviewed, Crossref-registered, and free to read under CC BY 4.0.

Submit a manuscript →Read our Scientific Writing policy →