Related Experiment Video
Updated: Jun 16, 2026

09:00
Advancing Dyslexia Assessment in Children Through Computerized Testing
Published on: August 16, 2024
Evaluating diagnostic tests: The area under the ROC curve and the balance of errors
1Imperial College, London. d.j.hand@imperial.ac.uk
Statistics in Medicine
|January 21, 2010
Summary
Evaluating diagnostic test accuracy is crucial for effective medicine. The area under the receiver operating characteristic (ROC) curve (AUC) has limitations in balancing misdiagnoses, prompting exploration of alternatives like the H measure.
Area of Science:
- Medical diagnostics
- Biostatistics
- Health informatics
Background:
- Accurate medical diagnosis is fundamental to patient care.
- Evaluating the performance of diagnostic tests is essential.
- The area under the receiver operating characteristic (ROC) curve (AUC) is a common accuracy measure.
Purpose of the Study:
- To explore the fundamental weaknesses of the AUC measure, particularly its inability to effectively balance different types of misdiagnoses.
- To introduce and describe an alternative measure, the H measure, for evaluating diagnostic test accuracy.
Main Methods:
- Analysis of the inherent properties of the AUC measure.
- Exploration of scenarios where ROC curves cross, highlighting AUC's limitations.
- Description of the proposed H measure as an alternative to AUC.
Main Results:
- The AUC measure has a fundamental weakness in balancing different kinds of misdiagnoses.
- This limitation is an inherent property of the AUC's definition, not just an arbitrary choice.
- The H measure is presented as a potential alternative that addresses this weakness.
Conclusions:
- The AUC measure, despite its widespread use, has significant limitations in accurately reflecting diagnostic test performance due to its handling of misdiagnoses.
- The proposed H measure offers a more effective approach to balancing different types of diagnostic errors.
- Further investigation into the H measure is warranted for improved diagnostic test evaluation.
Related Concept Videos
Receiver Operating Characteristic Plot
A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
Sensitivity, Specificity, and Predicted Value
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
Accuracy and Errors in Hypothesis Testing
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Errors In Hypothesis Tests
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
Detection of Gross Error: The Q Test
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
Critical Region, Critical Values and Significance Level
The critical region, critical value, and significance level are interdependent concepts crucial in hypothesis testing.
In hypothesis testing, a sample statistic is converted to a test statistic using z, t, or chi-square distribution. A critical region is an area under the curve in probability distributions demarcated by the critical value. When the test statistic falls in this region, it suggests that the null hypothesis must be rejected. As this region contains all those values of the test...
In hypothesis testing, a sample statistic is converted to a test statistic using z, t, or chi-square distribution. A critical region is an area under the curve in probability distributions demarcated by the critical value. When the test statistic falls in this region, it suggests that the null hypothesis must be rejected. As this region contains all those values of the test...
