A diagnostic accuracy study compares the test being evaluated, the index test, with a reference standard in the same participants. Report it with STARD 2015: describe both tests and their positivity thresholds, say what each reader knew, show participant flow in a diagram, give the 2x2 table and state sensitivity and specificity with 95% confidence intervals.
STARD 2015 (Standards for Reporting Diagnostic Accuracy studies) is a checklist of 30 essential items. It was published in October 2015 in three journals at once (BMJ, Radiology and Clinical Chemistry) and replaces the original 2003 STARD. The checklist has 34 rows because items 10, 12, 13 and 21 are split into a and b parts. The checklist and flow diagram template are on the EQUATOR Network's STARD page; the STARD 2015 full text is free.
Choosing STARD is covered by our guide to CONSORT, PRISMA, STROBE and other reporting guidelines and the guideline finder for surgical and clinical authors. This post starts where they stop: what each item asks you to write.
Define the index test, reference standard and target condition of your diagnostic accuracy study
STARD 2015 defines three terms the paper rests on. The index test is the test under evaluation. The reference standard is the best available method for establishing whether the target condition is present or absent. The target condition is the disease or condition the index test is expected to detect. STARD 2015 adds that a gold standard would be an error-free reference standard.
Describe both tests well enough for another team to repeat them, and say why you chose the reference standard. State the positivity cut-off for each test and whether it was fixed before the study or chosen from the data. A cut-off picked because it gave the best result is exploratory, and readers need to be told.
Describe who was tested: STARD asks whether participants formed a consecutive, random or convenience series, and on what basis they were identified. This matters because STARD 2015 notes that the relative numbers of false positive and false negative results vary across settings, depending on how patients present and which tests they have already had.
Say what each reader knew and how long separated the two tests
Item 13a covers the index test: state whether the people who performed or read it had access to the reference standard result and to clinical information about the participants. Item 13b asks the reverse for those assessing the reference standard. The STARD 2015 explanation and elaboration paper says readers need this information to interpret the findings, so report it even when nobody was blinded.
Item 22 asks for the interval between the two tests and any clinical interventions in between, because the target condition can change during a delay.
Report inconclusive and missing results instead of dropping them
Item 15 asks how indeterminate results on either test were handled, and item 16 asks the same for missing data. Under item 15, the explanation and elaboration paper encourages authors always to report how often indeterminate results occurred, with reasons, and any failures to complete testing. It warns that ignoring indeterminate results can bias accuracy estimates if they do not occur at random. Options include excluding them, reporting them as a separate category, or recalculating accuracy under worst-case and best-case assumptions.
A STARD 2015 checklist, item by item
Copy this table into your draft and add the page where each item appears. The wording is condensed, so check it against the official checklist. Registration is not only a trial item: items 28 to 30 ask for the registration number and registry, where the protocol can be read, and funding with the funder's role. If the study was not registered, say so rather than leaving the item blank.
| Section | Item | What to report |
|---|---|---|
| Title or abstract | 1 | Identify it as a diagnostic accuracy study, naming an accuracy measure |
| Abstract | 2 | Structured summary (see STARD for Abstracts) |
| Introduction | 3 | Background, including the intended use of the index test |
| Introduction | 4 | Objectives and hypotheses |
| Methods | 5 | Prospective or retrospective data collection |
| Methods | 6 | Eligibility criteria |
| Methods | 7 | How participants were identified, such as by symptoms |
| Methods | 8 | Setting, location and dates |
| Methods | 9 | Consecutive, random or convenience series |
| Methods | 10a | Index test, in enough detail to repeat it |
| Methods | 10b | Reference standard, in enough detail to repeat it |
| Methods | 11 | Why this reference standard, if alternatives exist |
| Methods | 12a | Index test cut-off or categories, with rationale; pre-specified or exploratory |
| Methods | 12b | The same for the reference standard |
| Methods | 13a | What index test readers knew of clinical information and reference results |
| Methods | 13b | What reference standard assessors knew of clinical information and index results |
| Methods | 14 | Methods for estimating or comparing accuracy |
| Methods | 15 | How indeterminate results on either test were handled |
| Methods | 16 | How missing data on either test were handled |
| Methods | 17 | Analyses of variability in accuracy, pre-specified or exploratory |
| Methods | 18 | Intended sample size and how it was determined |
| Results | 19 | Flow of participants, using a diagram |
| Results | 20 | Baseline demographic and clinical characteristics |
| Results | 21a | Severity of disease in those with the target condition |
| Results | 21b | Alternative diagnoses in those without it |
| Results | 22 | Interval and any interventions between the two tests |
| Results | 23 | Cross tabulation of the two tests' results |
| Results | 24 | Accuracy estimates and their precision, such as 95% CIs |
| Results | 25 | Adverse events from either test |
| Discussion | 26 | Limitations: sources of bias, statistical uncertainty, generalisability |
| Discussion | 27 | Implications for practice |
| Other information | 28 | Registration number and name of registry |
| Other information | 29 | Where the full study protocol can be accessed |
| Other information | 30 | Funding and the role of funders |
Worked example: the flow diagram and the 2x2 table
The study below is invented to illustrate the structure; its numbers are not real data. Adults attending an emergency department with suspected deep vein thrombosis had a new point-of-care blood test (the index test) and then compression ultrasound (the reference standard). The participant flow diagram (item 19) holds these counts:
- Assessed for eligibility: 480
- Excluded: 60 (41 did not meet the eligibility criteria; 19 declined)
- Received the index test: 420
- Index test inconclusive: 10 (device error); all 10 had the reference standard, which showed the target condition in 3 and not in 7
- Index test positive: 146; negative: 264
- No reference standard: 6 (left before ultrasound; 4 index positive, 2 index negative)
- Both tests conclusive: 404 (true positive 88, false positive 54, false negative 12, true negative 250)
Item 23 asks for the cross tabulation of the 404 complete results, which is the 2x2 table. The explanation and elaboration paper prefers actual numbers to percentages alone, noting that authors' mistakes in calculating sensitivity and specificity are not rare.
| Index test | Reference standard positive | Reference standard negative | Total |
|---|---|---|---|
| Positive | 88 (true positive) | 54 (false positive) | 142 |
| Negative | 12 (false negative) | 250 (true negative) | 262 |
| Total | 100 | 304 | 404 |
Report sensitivity and specificity with 95% confidence intervals
Sensitivity is the proportion of people with the target condition whom the index test calls positive. Specificity is the proportion without it whom the test calls negative. Item 24 asks for each estimate with its precision, such as a 95% confidence interval (CI). Name the interval method among your methods for estimating accuracy (item 14). How to word an interval is covered in reporting p-values, confidence intervals and effect sizes.
| Measure | Calculation | Estimate (95% CI, Wilson score) |
|---|---|---|
| Sensitivity | 88/100 | 88.0% (80.2 to 93.0) |
| Specificity | 250/304 | 82.2% (77.5 to 86.1) |
| Positive predictive value (PPV) | 88/142 | 62.0% (53.8 to 69.5) |
| Negative predictive value (NPV) | 250/262 | 95.4% (92.2 to 97.4) |
| Prevalence in this sample | 100/404 | 24.8% |
Counting the 10 inconclusive results as wrong gives the worst case: the 3 with the condition become false negatives and the 7 without it false positives. Sensitivity becomes 88/103 (85.4%) and specificity 250/311 (80.4%). Report both and say which is primary. Call these estimates, not the test's "true" values.
Of [n] participants who had both the index test and the reference standard, [n] ([%]) had [target condition]. The index test was positive in [n] of [n] participants with the condition (sensitivity [%], 95% CI [lower] to [upper]) and negative in [n] of [n] without it (specificity [%], 95% CI [lower] to [upper]). [n] index test results were inconclusive; counting them as incorrect gave a sensitivity of [%] and a specificity of [%]. Confidence intervals were calculated with the [method] method.
Predictive values from one diagnostic accuracy study do not transfer to another setting
PPV is the proportion of positive results that are true positives, and NPV the proportion of negative results that are true negatives. Both depend mathematically on prevalence: as prevalence rises, PPV rises and NPV falls.
Suppose, for illustration, the same sensitivity and specificity held where 5% of those tested have the condition. Per 1,000 people, 50 have it and 44 test positive; of the 950 without it, about 169 test positive and 781 negative. PPV falls from 62.0% to 44/213 (20.7%), and NPV rises to 781/787 (99.2%).
That assumption is itself shaky. Sensitivity and specificity are not mathematically tied to prevalence, but nor are they fixed properties of a test. A study of 23 meta-analyses found they often vary with prevalence, probably through differences in patient spectrum. Report predictive values for the population you studied, and describe that population's disease severity and alternative diagnoses (items 21a and 21b) so readers can judge how far it resembles theirs.
Abstracts, AI-based tests and prediction models have separate guidance
STARD for Abstracts, published in 2017, lists 11 essential items for the abstract, including accuracy estimates with 95% CIs and the registration number. STARD-AI, published in Nature Medicine in September 2025, extends STARD 2015 to diagnostic accuracy studies centred on artificial intelligence; its final checklist has 40 items. Developing or validating a score that combines several predictors is prediction-model research, reported with TRIPOD+AI (2024).
Measuring whether your department used a test in line with an agreed standard is a clinical audit, not an accuracy study; see how to publish a clinical audit or quality improvement project.
At Directive Publications, the author guidelines ask you to match your manuscript to the relevant EQUATOR checklist; for a diagnostic accuracy study, that is STARD 2015. If you find an error in this guide, report a problem to us.