ISSN-registered · Peer-reviewed · Open Access
JournalsAboutContact
Scientific Writing

Reporting a Diagnostic Accuracy Study with STARD 2015

DE By Directive Editorial Team, Directive Publications ·27 Sep 2026 ·7 min read
Reporting a Diagnostic Accuracy Study with STARD 2015

A diagnostic accuracy study compares the test being evaluated, the index test, with a reference standard in the same participants. Report it with STARD 2015: describe both tests and their positivity thresholds, say what each reader knew, show participant flow in a diagram, give the 2x2 table and state sensitivity and specificity with 95% confidence intervals.

STARD 2015 (Standards for Reporting Diagnostic Accuracy studies) is a checklist of 30 essential items. It was published in October 2015 in three journals at once (BMJ, Radiology and Clinical Chemistry) and replaces the original 2003 STARD. The checklist has 34 rows because items 10, 12, 13 and 21 are split into a and b parts. The checklist and flow diagram template are on the EQUATOR Network's STARD page; the STARD 2015 full text is free.

Choosing STARD is covered by our guide to CONSORT, PRISMA, STROBE and other reporting guidelines and the guideline finder for surgical and clinical authors. This post starts where they stop: what each item asks you to write.

Define the index test, reference standard and target condition of your diagnostic accuracy study

STARD 2015 defines three terms the paper rests on. The index test is the test under evaluation. The reference standard is the best available method for establishing whether the target condition is present or absent. The target condition is the disease or condition the index test is expected to detect. STARD 2015 adds that a gold standard would be an error-free reference standard.

Describe both tests well enough for another team to repeat them, and say why you chose the reference standard. State the positivity cut-off for each test and whether it was fixed before the study or chosen from the data. A cut-off picked because it gave the best result is exploratory, and readers need to be told.

Describe who was tested: STARD asks whether participants formed a consecutive, random or convenience series, and on what basis they were identified. This matters because STARD 2015 notes that the relative numbers of false positive and false negative results vary across settings, depending on how patients present and which tests they have already had.

Say what each reader knew and how long separated the two tests

Item 13a covers the index test: state whether the people who performed or read it had access to the reference standard result and to clinical information about the participants. Item 13b asks the reverse for those assessing the reference standard. The STARD 2015 explanation and elaboration paper says readers need this information to interpret the findings, so report it even when nobody was blinded.

Item 22 asks for the interval between the two tests and any clinical interventions in between, because the target condition can change during a delay.

Report inconclusive and missing results instead of dropping them

Item 15 asks how indeterminate results on either test were handled, and item 16 asks the same for missing data. Under item 15, the explanation and elaboration paper encourages authors always to report how often indeterminate results occurred, with reasons, and any failures to complete testing. It warns that ignoring indeterminate results can bias accuracy estimates if they do not occur at random. Options include excluding them, reporting them as a separate category, or recalculating accuracy under worst-case and best-case assumptions.

A STARD 2015 checklist, item by item

Copy this table into your draft and add the page where each item appears. The wording is condensed, so check it against the official checklist. Registration is not only a trial item: items 28 to 30 ask for the registration number and registry, where the protocol can be read, and funding with the funder's role. If the study was not registered, say so rather than leaving the item blank.

SectionItemWhat to report
Title or abstract1Identify it as a diagnostic accuracy study, naming an accuracy measure
Abstract2Structured summary (see STARD for Abstracts)
Introduction3Background, including the intended use of the index test
Introduction4Objectives and hypotheses
Methods5Prospective or retrospective data collection
Methods6Eligibility criteria
Methods7How participants were identified, such as by symptoms
Methods8Setting, location and dates
Methods9Consecutive, random or convenience series
Methods10aIndex test, in enough detail to repeat it
Methods10bReference standard, in enough detail to repeat it
Methods11Why this reference standard, if alternatives exist
Methods12aIndex test cut-off or categories, with rationale; pre-specified or exploratory
Methods12bThe same for the reference standard
Methods13aWhat index test readers knew of clinical information and reference results
Methods13bWhat reference standard assessors knew of clinical information and index results
Methods14Methods for estimating or comparing accuracy
Methods15How indeterminate results on either test were handled
Methods16How missing data on either test were handled
Methods17Analyses of variability in accuracy, pre-specified or exploratory
Methods18Intended sample size and how it was determined
Results19Flow of participants, using a diagram
Results20Baseline demographic and clinical characteristics
Results21aSeverity of disease in those with the target condition
Results21bAlternative diagnoses in those without it
Results22Interval and any interventions between the two tests
Results23Cross tabulation of the two tests' results
Results24Accuracy estimates and their precision, such as 95% CIs
Results25Adverse events from either test
Discussion26Limitations: sources of bias, statistical uncertainty, generalisability
Discussion27Implications for practice
Other information28Registration number and name of registry
Other information29Where the full study protocol can be accessed
Other information30Funding and the role of funders

Worked example: the flow diagram and the 2x2 table

The study below is invented to illustrate the structure; its numbers are not real data. Adults attending an emergency department with suspected deep vein thrombosis had a new point-of-care blood test (the index test) and then compression ultrasound (the reference standard). The participant flow diagram (item 19) holds these counts:

  1. Assessed for eligibility: 480
  2. Excluded: 60 (41 did not meet the eligibility criteria; 19 declined)
  3. Received the index test: 420
  4. Index test inconclusive: 10 (device error); all 10 had the reference standard, which showed the target condition in 3 and not in 7
  5. Index test positive: 146; negative: 264
  6. No reference standard: 6 (left before ultrasound; 4 index positive, 2 index negative)
  7. Both tests conclusive: 404 (true positive 88, false positive 54, false negative 12, true negative 250)

Item 23 asks for the cross tabulation of the 404 complete results, which is the 2x2 table. The explanation and elaboration paper prefers actual numbers to percentages alone, noting that authors' mistakes in calculating sensitivity and specificity are not rare.

Index testReference standard positiveReference standard negativeTotal
Positive88 (true positive)54 (false positive)142
Negative12 (false negative)250 (true negative)262
Total100304404

Report sensitivity and specificity with 95% confidence intervals

Sensitivity is the proportion of people with the target condition whom the index test calls positive. Specificity is the proportion without it whom the test calls negative. Item 24 asks for each estimate with its precision, such as a 95% confidence interval (CI). Name the interval method among your methods for estimating accuracy (item 14). How to word an interval is covered in reporting p-values, confidence intervals and effect sizes.

MeasureCalculationEstimate (95% CI, Wilson score)
Sensitivity88/10088.0% (80.2 to 93.0)
Specificity250/30482.2% (77.5 to 86.1)
Positive predictive value (PPV)88/14262.0% (53.8 to 69.5)
Negative predictive value (NPV)250/26295.4% (92.2 to 97.4)
Prevalence in this sample100/40424.8%

Counting the 10 inconclusive results as wrong gives the worst case: the 3 with the condition become false negatives and the 7 without it false positives. Sensitivity becomes 88/103 (85.4%) and specificity 250/311 (80.4%). Report both and say which is primary. Call these estimates, not the test's "true" values.

Of [n] participants who had both the index test and the reference standard, [n] ([%]) had [target condition]. The index test was positive in [n] of [n] participants with the condition (sensitivity [%], 95% CI [lower] to [upper]) and negative in [n] of [n] without it (specificity [%], 95% CI [lower] to [upper]). [n] index test results were inconclusive; counting them as incorrect gave a sensitivity of [%] and a specificity of [%]. Confidence intervals were calculated with the [method] method.

Predictive values from one diagnostic accuracy study do not transfer to another setting

PPV is the proportion of positive results that are true positives, and NPV the proportion of negative results that are true negatives. Both depend mathematically on prevalence: as prevalence rises, PPV rises and NPV falls.

Suppose, for illustration, the same sensitivity and specificity held where 5% of those tested have the condition. Per 1,000 people, 50 have it and 44 test positive; of the 950 without it, about 169 test positive and 781 negative. PPV falls from 62.0% to 44/213 (20.7%), and NPV rises to 781/787 (99.2%).

That assumption is itself shaky. Sensitivity and specificity are not mathematically tied to prevalence, but nor are they fixed properties of a test. A study of 23 meta-analyses found they often vary with prevalence, probably through differences in patient spectrum. Report predictive values for the population you studied, and describe that population's disease severity and alternative diagnoses (items 21a and 21b) so readers can judge how far it resembles theirs.

Abstracts, AI-based tests and prediction models have separate guidance

STARD for Abstracts, published in 2017, lists 11 essential items for the abstract, including accuracy estimates with 95% CIs and the registration number. STARD-AI, published in Nature Medicine in September 2025, extends STARD 2015 to diagnostic accuracy studies centred on artificial intelligence; its final checklist has 40 items. Developing or validating a score that combines several predictors is prediction-model research, reported with TRIPOD+AI (2024).

Measuring whether your department used a test in line with an agreed standard is a clinical audit, not an accuracy study; see how to publish a clinical audit or quality improvement project.

At Directive Publications, the author guidelines ask you to match your manuscript to the relevant EQUATOR checklist; for a diagnostic accuracy study, that is STARD 2015. If you find an error in this guide, report a problem to us.

Frequently Asked Questions

What is the difference between sensitivity and positive predictive value?
Sensitivity is the proportion of people who have the target condition whom the index test correctly calls positive. Positive predictive value is the proportion of people with a positive result who actually have the condition. Positive predictive value rises and falls with prevalence; sensitivity is not mathematically tied to prevalence but can still vary between settings.
Does a diagnostic accuracy study need a flow diagram if every participant had both tests?
Yes. STARD 2015 item 19 asks for the flow of participants using a diagram, whatever the attrition. The diagram shows how many were assessed for eligibility, how many received each test and the numbers of true and false positive and negative results, so readers can check the arithmetic.
Should inconclusive index test results be left out of the 2x2 table?
They can be kept out of the main 2x2 table, but not out of the report. The STARD 2015 explanation and elaboration paper encourages authors always to report how many there were, with reasons, and warns that ignoring indeterminate results can bias accuracy estimates if they do not occur at random. A worst-case analysis that counts them as wrong shows readers how much they matter.
Is there a STARD checklist for AI-based diagnostic tests?
Yes. STARD-AI, published in Nature Medicine in September 2025, extends STARD 2015 to diagnostic accuracy studies centred on artificial intelligence. Its final checklist has 40 items, including modified STARD items and new AI-specific ones. The EQUATOR Network lists it alongside STARD 2015.
DE
Directive Editorial Team
Directive Publications

The editorial team at Directive Publications — an international open-access publisher of peer-reviewed medical and scientific journals.

Publishing your research?

Directive Publications is an open-access publisher — every article peer-reviewed, Crossref-registered, and free to read under CC BY 4.0.

Submit a manuscript →Read our Scientific Writing policy →