Related Experiment Video
Updated: Dec 31, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
A new concordant partial AUC and partial c statistic for imbalanced data in the evaluation of machine learning
André M Carrington1, Paul W Fieguth2, Hammad Qazi3
1Ottawa Hospital Research Institute, Ottawa, K1H 8L6, Canada. acarrington@ohri.ca.
This study introduces new statistical tools to better evaluate machine learning models when dealing with imbalanced datasets, where one class of data is much rarer than another. Traditional evaluation methods often fail to provide a clear picture in these scenarios. The authors developed two new metrics, a partial area under the curve and a partial c statistic, which offer more accurate and interpretable results. These new measures maintain the beneficial properties of standard metrics while focusing on the most relevant parts of the performance curve. By testing these tools on breast cancer datasets, the researchers demonstrated that their approach provides a more reliable way to assess diagnostic performance. This work helps practitioners make better decisions when training models on skewed data.
Area of Science:
- Computational statistics and machine learning evaluation metrics
- Biostatistics and concordant partial AUC research within diagnostic medicine
Background:
Standard evaluation metrics often struggle to provide meaningful insights when datasets contain highly unequal class distributions. Prior research has shown that traditional receiver-operator characteristic curves can become misleading in these specific scenarios. That uncertainty drove the development of various alternative metrics designed to address performance measurement in skewed environments. No prior work had resolved the inherent limitations regarding the interpretability of these existing partial measures. Many current approaches fail to account for the full spectrum of information regarding actual negative cases. This gap motivated the search for more robust statistical frameworks that preserve the integrity of original performance indicators. Researchers have long sought methods that maintain the mathematical properties of established whole-curve metrics. This paper addresses these challenges by introducing refined statistical tools for evaluating predictive performance in imbalanced settings.
Purpose Of The Study:
The aim of this study is to derive and propose a new concordant partial area under the curve and a partial c statistic. These measures are intended to help researchers understand and explain specific segments of the receiver-operator characteristic plot. The authors seek to address the limitations inherent in traditional evaluation metrics when applied to imbalanced datasets. Many existing alternatives fail to provide a full interpretation because they neglect information regarding actual negative cases. This research focuses on providing foundational measures that maintain the beneficial properties of standard whole-curve metrics. The authors motivate this work by highlighting the need for more accurate assessment tools in diagnostic testing and classification tasks. They intend to provide a robust framework that functions effectively across various data types, with a specific emphasis on low prevalence scenarios. This study establishes a clear path for improving how practitioners evaluate the performance of their predictive algorithms.
Main Methods:
The research team employed a rigorous mathematical derivation to construct their new partial performance indicators. Their review approach involved testing these metrics against classic receiver-operator characteristic examples established in previous literature. The authors utilized two distinct breast cancer datasets to verify the practical utility of their proposed statistical framework. They performed a comparative analysis to ensure that their partial measures aligned with established whole-measure counterparts. The study design focused on validating the continuous and discrete versions of the proposed metrics. Researchers systematically compared these results to ensure mathematical equivalence across different data configurations. This methodology prioritized the maintenance of standard metric characteristics while isolating specific segments of the performance curve. The team concluded their assessment by providing a detailed interpretation of the results to illustrate the necessity of their new approach.
Main Results:
The study confirms that the new partial measures maintain the expected equalities with their existing whole-measure counterparts. The researchers demonstrated that their concordant partial area under the receiver-operator characteristic curve preserves the essential characteristics of the full area under the curve. They successfully derived the first partial c statistic specifically for receiver-operator characteristic plots to provide an unbiased interpretation. The validation process using the Wisconsin and Ljubljana breast cancer datasets showed that the partial measures sum to the total values where expected. These results highlight that the new metrics effectively address the limitations of previous alternatives. The authors observed that their approach remains robust even when applied to datasets with low prevalence. Their findings illustrate that the partial measures provide a more informative evaluation than traditional methods for imbalanced data. The analysis confirms the mathematical consistency between the continuous and discrete versions of the proposed statistical tools.
Conclusions:
The authors propose that their new partial measures successfully maintain the core characteristics of standard whole-curve metrics. These tools provide a consistent framework for interpreting specific segments of performance curves in diagnostic testing. The researchers confirm that their partial area under the curve and partial c statistic are mathematically equivalent. Their analysis demonstrates that these partial measures sum correctly to the total values of the original metrics. The study suggests that these methods offer a more unbiased interpretation of performance than previous alternatives. These findings imply that practitioners can now evaluate models with greater clarity when dealing with low prevalence data. The authors propose that these measures are applicable across diverse datasets beyond the specific examples tested here. This work provides a foundation for more nuanced assessment of machine learning algorithms in clinical and research applications.
Frequently Asked Questions
The researchers propose a concordant partial area under the receiver-operator characteristic curve and a partial c statistic. These metrics allow for focused evaluation of specific curve segments while maintaining the mathematical properties of the full area under the curve, unlike previous alternatives that ignore negative case information.
The authors utilize the Wisconsin and Ljubljana breast cancer datasets to validate their proposed measures. These real-life benchmarks, alongside classic examples from Fawcett, demonstrate the validity of the new metrics by confirming expected equalities between partial and whole-measure counterparts.
A partial c statistic is necessary because it provides an unbiased interpretation for specific portions of a receiver-operator characteristic plot. This tool is required to address the limitations of existing measures that fail to fully account for actual negative data points.
The researchers employ a continuous and discrete approach to derive their measures from the area under the curve and c statistic. This dual-methodology ensures that the partial metrics remain consistent with their whole-measure counterparts during the evaluation process.
The study measures the validity of the new tools by comparing them against existing whole measures. The researchers confirm that the partial measures sum to the total values of the original metrics, ensuring mathematical consistency and reliability in performance assessment.
The authors propose that these measures may be combined with other receiver-operator characteristic techniques in future work. They suggest that these tools could also be tested for their value in datasets characterized by high prevalence.
Related Concept Videos
Receiver Operating Characteristic Plot
One-Way ANOVA: Unequal Sample Sizes
Bioequivalence Data: Statistical Interpretation
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Comparing the Survival Analysis of Two or More Groups
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...

