Research Methods & Reporting: STARD 2015: An Updated List of Essential Items For Reporting Diagnostic Accuracy Studies
Research Methods & Reporting: STARD 2015: An Updated List of Essential Items For Reporting Diagnostic Accuracy Studies
1 2 3 4
Patrick M Bossuyt , Johannes B Reitsma , David E Bruns , Constantine A Gatsonis , Paul P
5 6 7 89 10 11
Glasziou , Les Irwig , Jeroen G Lijmer , David Moher , Drummond Rennie , Henrica C W de
12 13 14 15 16 17 18 19 20
Vet , Herbert Y Kressel , Nader Rifai , Robert M Golub , Douglas G Altman , Lotty Hooft ,
1 1 21
Daniël A Korevaar , Jérémie F Cohen , for the STARD Group
1
Department of Clinical Epidemiology, Biostatistics and Bioinformatics, Academic Medical Centre, University of Amsterdam, Amsterdam, the
Netherlands; 2Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, University of Utrecht, Utrecht, the Netherlands;
3
Department of Pathology, University of Virginia School of Medicine, Charlottesville, VA, USA; 4Center for Statistical Sciences, Brown University
School of Public Health, Providence, RI, USA; 5Centre for Research in Evidence-Based Practice, Faculty of Health Sciences and Medicine, Bond
University, Gold Coast, Queensland, Australia; 6Screening and Diagnostic Test Evaluation Program, School of Public Health, University of Sydney,
Sydney, New South Wales, Australia; 7Department of Psychiatry, Onze Lieve Vrouwe Gasthuis, Amsterdam, the Netherlands; 8Clinical Epidemiology
Program, Ottawa Hospital Research Institute, Ottawa, Canada; 9School of Epidemiology, Public Health and Preventive Medicine, University of
Ottawa, Ottawa, Canada; 10Peer Review Congress, Chicago, IL, USA; 11Philip R Lee Institute for Health Policy Studies, University of California, San
Francisco, CA, USA; 12Department of Epidemiology and Biostatistics, EMGO Institute for Health and Care Research, VU University Medical Center,
Amsterdam, the Netherlands; 13Department of Radiology, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA;
14
Radiology Editorial Office, Boston, MA, USA; 15Department of Laboratory Medicine, Boston Children’s Hospital, Harvard Medical School, Boston,
MA, USA; 16Clinical Chemistry Editorial Office, Washington, DC, USA; 17Division of General Internal Medicine and Geriatrics and Department of
Preventive Medicine, Northwestern University Feinberg School of Medicine, Chicago, IL, USA; 18JAMA Editorial Office, Chicago, IL, USA; 19Centre
for Statistics in Medicine, Nuffield Department of Orthopaedics, Rheumatology and Musculoskeletal Sciences, University of Oxford, Oxford, UK;
20
Dutch Cochrane Centre, Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, University of Utrecht, Utrecht,
the Netherlands; 21INSERM UMR 1153 and Department of Pediatrics, Necker Hospital, AP-HP, Paris Descartes University, Paris, France.
As researchers, we talk and write about our studies, not just elements of study methods are often poorly described and
because we are happy—or disappointed—with the findings, but sometimes completely omitted, making both critical appraisal
also to allow others to appreciate the validity of our methods, and replication difficult, if not impossible. Sometimes study
to enable our colleagues to replicate what we did, and to disclose results are selectively reported, and other times researchers
our findings to clinicians, other health care professionals, and cannot resist unwarranted optimism in interpretation of their
decision makers, all of whom rely on the results of strong findings.2-4 These practices limit the value of the research and
research to guide their actions. any downstream products or activities, such as systematic
Unfortunately, deficiencies in the reporting of research have reviews and clinical practice guidelines.
been highlighted in several areas of clinical medicine.1 Essential
Reports of studies of medical tests are no exception. A growing Quality and Transparency of Health Research (EQUATOR)
number of evaluations have identified deficiencies in the website at [Link]/reporting-guidelines/stard.
reporting of test accuracy studies.5 These are studies in which In short, we invited the 2003 STARD group members to
a test is evaluated against a clinical reference standard, or gold participate in the updating process, nominate new members,
standard; the results are typically reported as estimates of the and comment on the general scope of the update. Suggested
test’s sensitivity and specificity, which express how good the new members were contacted. As a result, the STARD group
test is in correctly identifying patients as having the target has now grown to 85 members that include researchers, editors,
condition. Other accuracy statistics can be used as well, such journalists, evidence synthesis professionals, funders, and other
as the area under the receiver operating characteristics (ROC) stakeholders.
curve or positive and negative predictive values. STARD group members were then asked to suggest, and later
Despite their apparent simplicity, such studies are at risk of to endorse, proposed changes in a two round, web based survey.
bias.6 7 If not all patients undergoing testing are included in the This served to prepare a draft list of essential items, which was
final analysis, for example, or if only healthy controls are discussed in the steering committee in a two day meeting in
included, the estimates of test accuracy may not reflect the Amsterdam in September 2014. The list was then piloted in
performance of the test in clinical applications. Yet such crucial different groups: starting and advanced researchers, peer
information is often missing from study reports. reviewers, and editors.
It is now well established that sensitivity and specificity are not The general structure of STARD 2015 is similar to that of
fixed test properties. The relative number of false positive and STARD 2003. A one page document presents 30 items, grouped
false negative test results varies across settings, depending on under sections that follow the introduction, methods, results,
how patients present and which tests they have already and discussion (IMRAD) structure of a scientific article (see
undergone. Unfortunately, many authors also fail to completely table 1⇓). Several of the STARD 2015 items are identical to the
report the clinical context and when, where, and how they ones in the 2003 version. Others have been reworded, combined,
identified and recruited eligible study participants.8 In addition, or (if complex) split. A few have been added (see table 2⇓ for
sensitivity and specificity estimates can differ because of a summary of new items and table 3⇓ for key terms). A diagram
variable definitions of the reference standard against which the to describe the flow of participants through the study is now
test is being compared. Thus this information should be available expected in all reports (figure⇓).
in the study report.
The 2003 STARD statement Scope
To assist in the completeness and transparency of reporting STARD 2015 replaces the original version published in 2003;
diagnostic accuracy studies, a group of researchers, editors, and those who would like to refer to STARD are invited to cite this
other stakeholders developed a minimum list of essential items article. The list of essential items can be seen as a minimum set,
that should be included in every study report. The guiding and an informative study report will typically present more
principle for developing the list was to select items that, if information. Yet we hope to find all applicable items in a well
described, would help readers to judge the potential for bias in prepared report of a diagnostic accuracy study.
the study and appraise the applicability of the study findings Authors are invited to use STARD when preparing their study
and the validity of the authors’ conclusions and reports. Reviewers can use the list to verify that all essential
recommendations. information is available in a submitted manuscript and suggest
The resulting Standards for Reporting Diagnostic Accuracy changes if key items are missing.
(STARD) statement appeared in 2003 in two dozen journals.9 We trust that journals that endorsed STARD in 2003 or later
It was accompanied by editorials and commentaries in several will recommend the use of this updated version and encourage
other publications and endorsed by many more. compliance in submitted manuscripts. We hope that even more
Since the publication of STARD, several evaluations have journals, and journal organizations, will promote the use of this
pointed to small but statistically significant improvements in and comparable reporting guidelines. Funders and research
reporting accuracy studies (mean gain 1.4 items (95% institutions may promote or mandate adherence to STARD as
confidence interval 0.7 to 2.2)).5 10 Gradually, more of the a way to maximize the value of research and downstream
essential items are being reported, but the situation remains far products or activities.
from optimal. STARD may also be beneficial for reporting other studies that
evaluate the performance of tests. This includes prognostic
Methods for developing STARD 2015 studies, which can classify patients on the basis of whether a
future event happens; monitoring studies, in which tests are
The STARD steering committee periodically reviews the supposed to detect or predict an adverse event or lack of
literature for potentially relevant studies to inform a possible response; studies evaluating treatment selection markers; and
update. In 2013, the steering committee decided that the time more. We and others have found most of the STARD items
was right to update the checklist. useful when reporting and examining such studies, although
Updating had two major goals: first, to incorporate recent STARD primarily targets diagnostic accuracy studies.
evidence about sources of bias, applicability concerns, and Diagnostic accuracy is not the only expression of test
factors facilitating generous interpretation in test accuracy performance, nor is it always the most meaningful.12 Incremental
research, and, second, to make the list easier to use. In making accuracy from combining tests, relative to a single test, can be
modifications, we also considered harmonization with other more informative, for example.13 For continuous tests,
reporting guidelines, such as Consolidated Standards of dichotomization into test positives and negatives may not always
Reporting Trials (CONSORT) 2010.11 be indicated. In such cases, the desirable computational and
A complete description of the updating process and the graphical methods for expressing test performance are different,
justification for the changes are available on the Enhancing the although many of the methodological precautions would be the
same, and STARD can help in reporting the study in an Increasing value, reducing waste
informative way. Other reporting guidelines target more specific
forms of tests, such as Transparent Reporting of a Multivariable The STARD steering committee is aware that building a list of
Prediction Model for Individual Prognosis or Diagnosis essential items is not sufficient to achieve substantial
(TRIPOD) for multivariable prediction models.14 improvements in reporting completeness, as the modest
improvement after introduction of the 2003 list has shown. We
Although STARD focuses on full study reports of test accuracy
see this list not as the final product, but as the starting point for
studies, the items can also be helpful when writing conference
building more specific instruments to stimulate complete and
abstracts, including information in trial registries, and
transparent reporting, such as a checklist and a writing aid for
developing protocols for such studies. Additional initiatives are
authors, tools for reviewers and editors, instruction videos, and
underway to provide more specific guidance for each of these
teaching materials, all based on this STARD list of essential
applications.
items.
Incomplete reporting has been identified as one of the sources
STARD extensions and applications of avoidable waste in biomedical research.1 Since STARD was
The STARD statement was designed to apply to all types of initiated, several other initiatives have been undertaken to
medical tests. The STARD group believed that a single checklist, enhance the reproducibility of research and promote greater
for all diagnostic accuracy studies, would be more widely transparency.20 Multiple factors are at stake, but incomplete
disseminated and more easily accepted by authors, peer reporting is one of them. We hope that this update of STARD,
reviewers, and journal editors than separate lists for different together with additional implementation initiatives, will help
types of tests such as imaging, biochemistry, or histopathology. authors, editors, reviewers, readers, and decision makers to
Having a general list may necessitate additional instructions for collect, appraise, and apply the evidence needed to strengthen
informative reporting, with more information for specific types decisions and recommendations about medical tests. In the end,
of tests, specific applications, or specific forms of analysis. Such we are all to benefit from more informative and transparent
guidance could describe the preferred methods for studying and reporting: as researchers, as healthcare professionals, as payers,
reporting measurement uncertainty, for example, without and as patients.
changing any of the other STARD items. The STARD group
welcomes the development of such STARD extensions and This article is being simultaneously published in October 2015 by The
invites interested groups to contact the STARD executive BMJ, Radiology, and Clinical Chemistry. This article is published under
committee before developing them. the Creative Commons CC BY-NC license [Link]
licenses/by-nc/4.0.
Other groups may want to develop additional guidance to
STARD Group collaborators: Todd Alonzo, Douglas G Altman, Augusto
facilitate the use of STARD for specific applications. An
Azuara-Blanco, Lucas Bachmann, Jeffrey Blume, Patrick M Bossuyt,
example of such a STARD application was prepared for history
Isabelle Boutron, David Bruns, Harry Büller, Frank Buntinx, Sarah Byron,
taking and physical examination.15 Another type of application
Stephanie Chang, Jérémie F Cohen, Richelle Cooper, Joris de Groot,
is the use of STARD for specific target conditions such as
Henrica C W de Vet, Jon Deeks, Nandini Dendukuri, Jac Dinnes,
dementia.16
Kenneth Fleming, Constantine A Gatsonis, Paul P Glasziou, Robert M
Golub, Gordon Guyatt, Carl Heneghan, Jørgen Hilden, Lotty Hooft, Rita
Availability Horvath, Myriam Hunink, Chris Hyde, John Ioannidis, Les Irwig, Holly
The new STARD 2015 list and all related documents can be Janes, Jos Kleijnen, André Knottnerus, Daniël A Korevaar, Herbert Y
found on the STARD pages of the EQUATOR website. Kressel, Stefan Lange, Mariska Leeflang, Jeroen G Lijmer, Sally Lord,
EQUATOR is an international initiative that seeks to improve Blanca Lumbreras, Petra Macaskill, Erik Magid, Susan Mallett, Matthew
the value of published health research literature by promoting McInnes, Barbara McNeil, Matthew McQueen, David Moher, Karel
transparent and accurate reporting and wider use of robust Moons, Katie Morris, Reem Mustafa, Nancy Obuchowski, Eleanor
reporting guidelines.17 18 The STARD group believes that Ochodo, Andrew Onderdonk, John Overbeke, Nitika Pai, Rosanna
working more closely with EQUATOR and other reporting Peeling, Margaret Pepe, Steffen Petersen, Christopher Price, Philippe
guideline developers will help us to better reach shared Ravaud, Johannes B. Reitsma, Drummond Rennie, Nader Rifai, Anne
objectives. We have updated the 2003 explanation and Rutjes, Holger Schunemann, David Simel, Iveta Simera, Nynke Smidt,
elaboration document, which can also be found at the Ewout Steyerberg, Sharon Straus, William Summerskill, Yemisi
EQUATOR website. This document explains the rationale for Takwoingi, Matthew Thompson, Ann van de Bruel, Hans van Maanen,
each item and gives examples. Andrew Vickers, Gianni Virgili, Stephen Walter, Wim Weber, Marie
Westwood, Penny Whiting, Nancy Wilczynski, Andreas Ziegler.
The STARD list is released under a Creative Commons license.
This allows everyone to use and distribute the work if they Contributors: All authors confirm they have contributed to the intellectual
acknowledge the source. The STARD statement was originally content of this paper and have met the following 3 requirements: (a)
reported in English, but several groups have worked on significant contributions to the conception and design, acquisition of
translations in other languages. We welcome such translations, data, or analysis and interpretation of data; (b) drafting or revising the
which are preferably developed by groups of researchers, by article for intellectual content; and (c) final approval of the published
use of a cyclical development process, with back-translation to article.
the original language and user testing.19 We have also applied Funding: There was no explicit funding for the development of STARD
for a trademark for STARD to ensure that the steering committee 2015. The Academic Medical Center of the University of Amsterdam,
has the exclusive right to use the word “STARD” to identify the Netherlands, partly funded the meeting of the STARD steering group
goods or services. but had no influence on the development or dissemination of the list of
essential items. STARD steering group members and STARD group
members covered additional personal costs individually.
Competing interests: All authors have completed the Clinical Chemistry 12 Bossuyt PM, Reitsma JB, Linnet K, Moons KG. Beyond diagnostic accuracy: the clinical
utility of diagnostic tests. Clin Chem 2012;58:1636-43.
author disclosure form: N Rifai works for Clinical Chemistry, AACC; C 13 Moons KG, de Groot JA, Linnet K, Reitsma JB, Bossuyt PM. Quantifying the added value
A Gatsonisis a member of RSNA Research Development Committee. of a diagnostic test or marker. Clin Chem 2012;58:1408-17.
14 Collins GS, Reitsma JB, Altman DG, Moons KG. Transparent reporting of a multivariable
prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement.
1 Glasziou P, Altman DG, Bossuyt P, et al. Reducing waste from incomplete or unusable BMJ 2015;350:g7594.
reports of biomedical research. Lancet 2014;383:267-76. 15 Simel DL, Rennie D, Bossuyt PM. The STARD statement for reporting diagnostic accuracy
2 Boutron I, Dutton S, Ravaud P, Altman DG. Reporting and interpretation of randomized studies: application to the history and physical examination. J Gen Intern Med
controlled trials with statistically nonsignificant results for primary outcomes. JAMA 2008;23:768-74.
2010;303:2058-64. 16 Noel-Storr AH, McCleery JM, Richard E, et al. Reporting standards for studies of diagnostic
3 Ochodo EA, de Haan MC, Reitsma JB, et al. Overinterpretation and misreporting of test accuracy in dementia: the STARDdem Initiative. Neurology 2014;83:364-73.
diagnostic accuracy studies: evidence of “spin.” Radiology 2013;267:581-8. 17 Altman DG, Simera I, Hoey J, Moher D, Schulz K. EQUATOR: reporting guidelines for
4 Mathieu S, Boutron I, Moher D, Altman DG, Ravaud P. Comparison of registered and health research. Lancet 2008;371:1149-50.
published primary outcomes in randomized controlled trials. JAMA 2009;302:977-84. 18 Simera I, Moher D, Hirst A, et al. Transparent and accurate reporting increases reliability,
5 Korevaar DA, Wang J, van Enst WA, et al. Reporting diagnostic accuracy studies: some utility, and impact of your research: reporting guidelines and the EQUATOR Network.
improvements after 10 years of STARD. Radiology 2015;274:781-9. BMC Med 2010;8:24.
6 Lijmer JG, Mol BW, Heisterkamp S, et al. Empirical evidence of design-related bias in 19 Beaton DE, Bombardier C, Guillemin F, Ferraz MB. Guidelines for the process of
studies of diagnostic tests. JAMA 1999;282:1061-6. cross-cultural adaptation of self-report measures. Spine 2000;25:3186-91.
7 Whiting PF, Rutjes AW, Westwood ME, Mallett S. A systematic review classifies sources 20 Collins FS, Tabak LA. Policy: NIH plans to enhance reproducibility. Nature 2014;505:612-3.
of bias and variation in diagnostic test accuracy studies. J Clin Epidemiol
2013;66:1093-104. Accepted: 18 September 2015
8 Irwig L, Bossuyt P, Glasziou P, Gatsonis C, Lijmer J. Designing studies to ensure that
estimates of test accuracy are transferable. BMJ 2002;324:669-71.
9 Bossuyt PM, Reitsma JB, Bruns DE, et al. Towards complete and accurate reporting of Cite this as: BMJ 2015;351:h5527
studies of diagnostic accuracy: the STARD Initiative. Radiology 2003;226:24-8.
© Bossuyt et al 2015
10 Korevaar DA, van Enst WA, Spijker R, Bossuyt PM, Hooft L. Reporting quality of diagnostic
accuracy studies: a systematic review and meta-analysis of investigations on adherence
This is an Open Access article distributed in accordance with the terms of the Creative
to STARD. Evid Based Med 2014;19:47-54. Commons Attribution (CC BY 4.0) license, which permits others to distribute, remix, adapt
11 Schulz KF, Altman DG, Moher D, Group C. CONSORT 2010 statement: updated guidelines and build upon this work, for commercial use, provided the original work is properly cited.
for reporting parallel group randomised trials. J Clin Epidemiol 2010;63:834-40. See: [Link]
Tables
Table 1 (continued)
*At the start of each item row, authors should specify the page number of the manuscript where the item can be found.
No Item Rationale
2 Structured abstract Abstracts are increasingly used to identify key elements of study design and results.
3 Intended use and clinical role of the test Describing the targeted application of the test helps readers to interpret the implications of reported accuracy
estimates.
4 Study hypotheses Not having a specific study hypothesis may invite generous interpretation of the study results and “spin” in the
conclusions.
18 Sample size Readers want to appreciate the anticipated precision and power of the study and whether authors were successful
in recruiting the targeted number of participants.
26-27 Structured discussion To prevent jumping to unwarranted conclusions, authors are invited to discuss study limitations and draw conclusions
keeping in mind the targeted application of the evaluated tests (see item 3).
28 Registration Prospective test accuracy studies are trials, and, as such, they can be registered in clinical trial registries, such
as [Link], before their initiation, facilitating identification of their existence and preventing selective
reporting.
29 Protocol The full study protocol, with more information about the predefined study methods, may be available elsewhere,
to allow more fine grained critical appraisal.
30 Sources of funding Awareness of the potentially compromising effects of conflicts of interest between researchers’ obligations to abide
by scientific and ethical principles and other goals, such as financial ones; test accuracy studies are no exception.
Term Explanation
Medical test Any method for collecting additional information about the current or future health status of a patient
Index test The test under evaluation
Target condition The disease or condition that the index test is expected to detect
Clinical reference standard The best available method for establishing the presence or absence of the target condition; a gold standard would be an error-free
reference standard
Sensitivity Proportion of those with the target condition who test positive with the index test
Specificity Proportion of those without the target condition who test negative with the index test
Intended use of the test Whether the index test is used for diagnosis, screening, staging, monitoring, surveillance, prediction, prognosis, or other reasons
Role of the test The position of the index test relative to other tests for the same condition (for example, triage, replacement, add-on, new test)
Figure