ISSN-registered · Peer-reviewed · Open Access
JournalsAboutContact
Scientific Writing

Reporting a Clinical Prediction Model Using TRIPOD+AI

DE By Directive Editorial Team, Directive Publications ·4 Oct 2026 ·7 min read
Reporting a Clinical Prediction Model Using TRIPOD+AI

TRIPOD+AI is the 2024 reporting guideline for studies that develop, validate or update a clinical prediction model, whether built with regression or machine learning. Its checklist has 27 main items. Key items cover the study type, discrimination with calibration, how optimism was handled, the sample size, and the full model or its access restrictions.

TRIPOD stands for Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis. Collins and colleagues published TRIPOD+AI in The BMJ (British Medical Journal) in April 2024 (BMJ 2024;385:e078378), and the EQUATOR Network lists it. It replaces the 2015 TRIPOD checklist, which "should no longer be used", so where our guideline finder for surgical and clinical authors says TRIPOD, use TRIPOD+AI.

Risk scores, prediction rules, nomograms and algorithms that combine predictors to estimate an individual's risk fall under TRIPOD+AI. A study comparing one test with a reference standard is instead a diagnostic accuracy study, reported with STARD or STARD-AI. Prediction is not aetiology: the 2015 TRIPOD statement was not intended for aetiological studies, and a prediction model's coefficients should not be read as causal effects (see confounding and bias in observational studies).

Name the study type in the title and objectives

Items 1 and 4 ask you to state the study type. The TRIPOD scope names four:

  • Development. A new model, whose performance must still be evaluated, at least by internal validation.
  • External validation. An existing model tested in participants not used to develop it, for example from a later period (temporal) or another hospital or country (geographic).
  • Updating. An existing model recalibrated, extended or refitted.
  • A combination. For example, development plus external validation; every checklist item then applies.

Never call a model simply "validated": the TRIPOD+AI paper says "there is no such thing as a validated prediction model".

Report discrimination and calibration, not the AUC alone

Discrimination is how well the model separates people with the outcome from those without it. For a binary outcome it is the c statistic, which equals the area under the receiver operating characteristic curve (AUC). For time-to-event outcomes, use a censoring-aware version such as Harrell's or Uno's C. No AUC value is "good" everywhere; it depends on context and case mix.

Calibration is the agreement between predicted risks and observed outcomes. A model can rank patients well yet overestimate everyone's risk, so the AUC alone cannot show whether its risks can be trusted. The expanded TRIPOD+AI checklist expects both, with calibration plots, as a minimum. Riley and colleagues' guide to external validation asks for:

  • a smoothed calibration curve, with the distribution of predicted risks beneath it;
  • the calibration slope (ideal value 1; below 1 suggests predictions that are too extreme);
  • calibration-in-the-large (ideal value 0) and, for binary or time-to-event outcomes, the observed/expected (O/E) ratio (ideal value 1).

Give every estimate with a confidence interval (CI), overall and in key subgroups. Avoid the Hosmer–Lemeshow test, which depends on arbitrary grouping. If the model will guide decisions, add net benefit on a decision curve, at thresholds chosen in advance.

Show how you handled optimism and overfitting

Performance measured in the development data is optimistic, because the model partly fits noise, and more so in small samples. Item 12c asks for your internal validation method, including whether every modelling step, such as hyperparameter tuning, was repeated within it.

  • Bootstrapping. Rerun the whole process, including predictor selection and imputation, in each bootstrap sample. Each bootstrap model's optimism is its performance in its bootstrap sample minus that in the original data; subtract the average optimism from the apparent performance. Collins and colleagues' guide to model evaluation recommends at least 500 bootstraps.
  • k-fold cross-validation. Often comparable to bootstrapping.
  • Shrinkage or penalisation. Ridge or lasso regression, for example, pulls predictor effects towards zero to reduce overfitting, but a penalised model still needs internal validation.

The same guide says a random split into development and test sets is "generally advised against", because it discards data. Its halves are not independent, and it is not external validation.

Justify the sample size beyond events per variable

Item 10 asks how the sample size was reached for development and evaluation, and why it was enough, even if you used all available data. Riley and colleagues (2019) say rules of thumb such as 10 events per predictor parameter should be avoided. For binary and time-to-event outcomes, their criteria target:

  • a global shrinkage factor of 0.9 or more;
  • a difference of 0.05 or less between apparent and adjusted Nagelkerke R²;
  • precise estimation of the overall risk.

For external validation, a guide to validation sample size treats 100 events and 100 non-events only as a starting point, preferring a precision-based calculation. Why events limit regression, and how to tell a reviewer a small sample cannot support a model, are covered in choosing a statistical test and decoding statistical reviewer comments.

Handle missing predictors in development and at the point of use

Report the amount missing for each candidate predictor and the handling method, as in the Methods of a retrospective chart review, and the assumed reasons (item 11). Two points are model-specific:

  • confirmation that any imputation was done separately in training and test data, so test data cannot leak into the model;
  • item 27a: how a missing predictor value should be handled when the model is used in practice.

Report the full model so others can test it

Item 22 asks for the full model, as a formula, code, a software object or an application programming interface (API), so others can calculate a new person's risk and evaluate it. For regression, the 2015 checklist spelled this out: all coefficients and the intercept, or baseline survival at a stated time point.

Explain how to use the model, ideally with a worked calculation. This example is invented to illustrate the structure; its numbers are not real data. In a logistic model, the linear predictor (LP) is −5.20 + 0.04 × age in years + 0.85 × diabetes (1 = yes, 0 = no), and the predicted risk is 1 / (1 + exp(−LP)). For a 60-year-old with diabetes, LP = −1.95, so the predicted risk is 12.5%.

If the model cannot be shared, for example for commercial reasons, say so and give the access conditions. Analysis code is a separate item (18f); depositing it is covered in preparing supplementary materials.

A section-by-section TRIPOD+AI checklist

Compared with 2015, TRIPOD+AI adds patient and public involvement (item 19), open science (item 18) and fairness items that run from the background (3c) through subgroup performance (23a) to the discussion (25, 26). Item 13 asks why and how you corrected any class imbalance and how you recalibrated, because such corrections can make predicted risks too high. Studies of large language models (LLMs) have their own extension, TRIPOD-LLM.

Copy the table and add a "Page or line" column, as the TRIPOD+AI paper recommends; if an item does not apply, say why. Check the condensed wording against the official TRIPOD+AI checklist. D means development only; E, evaluation only.

SectionItemWhat to report
Title1Development or evaluation; population; outcome
Abstract2TRIPOD+AI for Abstracts
Introduction3a–4Context; existing models; intended users; health inequalities; objectives
Data5a–7Sources and dates per dataset; setting; eligibility; treatments; pre-processing
Outcome and predictors8a–9cDefinitions and timing; blinding; predictor choice (D)
Sample size, missing data10–11Justification per dataset; handling of missing data
Analysis12a–12gData use, predictor handling, model type, tuning, internal validation (D); performance measures; updating and how predictions were made (E)
Imbalance, fairness, output13–15Imbalance methods; fairness approaches; output and thresholds (D)
Data differences16Development versus evaluation data
Ethics17Committee; consent or waiver
Open science18a–18fFunding; conflicts of interest; protocol; registration; data and code sharing
Patient and public involvement19Details, or "no involvement"
Participants20a–21Flow; characteristics; comparison with development data (E); participants and events per analysis
Model specification22Full model; access restrictions (D)
Performance and updating23a–24Estimates with CIs, including key subgroups; updated model (E)
Discussion25–27cInterpretation and fairness; limitations; use in practice (D); next steps

Write the abstract with TRIPOD+AI for Abstracts

Item 2 points to TRIPOD+AI for Abstracts: 13 items, updated from the 2020 version, including performance estimates with CIs. Items 7 (model type, building steps and internal validation method) and 10 (the predictors in the final model) apply to development only.

Template sentences for the Methods and Results

Fill the brackets; delete what does not apply.

Methods. We [developed / externally validated / updated] a model to predict [outcome] within [time horizon] in [target population], reported following TRIPOD+AI. Data came from [source], [start date] to [end date]. The sample size ([n] participants, [n] events) was justified by [Riley et al.'s criteria / a precision-based calculation]. Missing values (Table [n]) were handled by [method] within each bootstrap sample. We fitted [model type] with [penalisation] and corrected for optimism with [number] bootstrap samples, repeating every modelling step. We assessed discrimination ([c statistic / Harrell's C]), calibration (smoothed plot, slope, calibration-in-the-large) and net benefit at thresholds of [range]. For external validation: we applied the published [equation / code] ([reference]) unchanged and compared development and validation data (Table [n]).

Results. Of [n] participants, [n] had [outcome]. The full equation of the [n]-predictor model is in [table / supplement], and the code is at [repository, DOI]. The [optimism-corrected / validation] c statistic was [value] (95% CI [lower] to [upper]), the calibration slope [value] and calibration-in-the-large [value] (CIs in Table [n]); Figure [n] shows the calibration plot. Subgroup performance is in Table [n].

Before submitting, check the paper against PROBAST+AI (2025), the updated Prediction model Risk Of Bias ASsessment Tool (PROBAST).

At Directive Publications, our author guidelines ask you to match your manuscript to the relevant EQUATOR Network checklist. The checklists they name do not include TRIPOD+AI, but it is the guideline EQUATOR lists for prediction models. Code sharing is strongly encouraged, not required; if yours cannot be shared, our data availability policy asks you to say so and explain why. If you spot an error, please report the problem to us.

Frequently Asked Questions

Does TRIPOD+AI apply to a logistic regression model, or only to machine learning?
TRIPOD+AI applies to both. Its authors describe it as harmonised guidance for prediction model studies, whether the model was built with regression or with machine learning. Because TRIPOD+AI supersedes the 2015 TRIPOD checklist, a conventional logistic or Cox model is now reported with TRIPOD+AI.
What is the difference between internal and external validation of a prediction model?
Internal validation estimates performance in the development data, by bootstrapping or cross-validation, to correct the optimism of the apparent performance. External validation tests the model in participants who were not used to develop it, such as patients from a later period or another hospital. Randomly splitting one dataset does not give external validation, and the two halves are not independent.
Which calibration measures should a prediction model paper report?
The expanded TRIPOD+AI checklist expects calibration, including calibration plots, as a minimum. Riley and colleagues' BMJ guide to external validation recommends a calibration plot with a smoothed calibration curve, the calibration slope (ideal value 1) and calibration-in-the-large (ideal value 0), and for binary or time-to-event outcomes the observed/expected ratio (ideal value 1). Give each measure with a confidence interval. Do not use the Hosmer-Lemeshow test, whose P value changes with the number of groups chosen and the sample size, and which does not show the size or direction of any miscalibration.
Do I have to publish my model's coefficients or code under TRIPOD+AI?
TRIPOD+AI item 22 asks for enough of the model for others to make predictions in new individuals and evaluate it: the equation for a regression model, or code, a software object or an application programming interface for a model that cannot be written as an equation. If the model cannot be made publicly available, for example for commercial reasons, state that clearly and give the conditions for access.
DE
Directive Editorial Team
Directive Publications

The editorial team at Directive Publications — an international open-access publisher of peer-reviewed medical and scientific journals.

Publishing your research?

Directive Publications is an open-access publisher — every article peer-reviewed, Crossref-registered, and free to read under CC BY 4.0.

Submit a manuscript →Read our Scientific Writing policy →