Related Experiment Videos
Multireader, multicase receiver operating characteristic analysis: an empirical comparison of five methods
Nancy A Obuchowski1, Sergey V Beiden, Kevin S Berbaum
1Departments of Biostatistics and Epidemiology, Cleveland Clinic Foundation, Cleveland, OH 44195, USA. nobuchow@bio.ri.ccf.org
Academic Radiology
|September 8, 2004
Summary
Comparing statistical methods for multireader, multicase (MRMC) receiver operating characteristic (ROC) studies reveals key differences. While several methods show concordance in fixed-reader models, variations emerge in random-reader models, impacting generalizability.
Area of Science:
- Medical Statistics
- Diagnostic Accuracy Research
- Biostatistics
Background:
- Multireader, multicase (MRMC) receiver operating characteristic (ROC) studies are crucial for evaluating diagnostic test performance.
- Various statistical methodologies exist for analyzing MRMC ROC data, each with unique assumptions and outputs.
- Understanding the concordance and discrepancies among these methods is essential for accurate interpretation of study results.
Purpose of the Study:
- To raise awareness of different statistical methods for MRMC ROC studies.
- To assess the concordance of results obtained from various MRMC methods using published datasets.
- To identify factors influencing the differences in outcomes across these analytical approaches.
Main Methods:
- Reanalysis of data from three previously published MRMC ROC studies.
- Application of five distinct MRMC statistical methods to the datasets.
- Reporting of 95% confidence intervals (CIs) for ROC areas, P values for accuracy comparisons, and CIs for mean differences in ROC areas for each method.
Main Results:
- Significant differences in P values and CIs were observed between parametric and nonparametric accuracy estimates, and between random-reader and fixed-reader models.
- For fixed-reader models, Dorfman-Berbaum-Metz (DBM), Obuchowski-Rockette, Beiden-Wagner-Campbell, and Song's multivariate Wilcoxon-Mann-Whitney (WMW) methods yielded highly similar results.
- In random-reader models, DBM, Obuchowski-Rockette, and Beiden-Wagner-Campbell showed comparable inferences, though Beiden-Wagner-Campbell produced broader CIs. Ishwaran's hierarchical ROC sometimes identified significance missed by others.
Conclusions:
- The choice and application of MRMC methods necessitate understanding the distinction between random-reader and fixed-reader models and their implications for generalizability.
- Awareness of the underlying assumptions of each MRMC method is critical for appropriate use.
- Limitations in reader variability within smaller studies (e.g., five or six readers) can influence the reliability and interpretation of MRMC analyses.