Related Experiment Video
Updated: Jun 26, 2025

Split Point Analysis and Uncertainty Quantification of Thermal-Optical Organic/Elemental Carbon Measurements
Published on: September 7, 2019
Commonly used software tools produce conflicting and overly-optimistic AUPRC values
Wenyu Chen1, Chen Miao1, Zhenghao Zhang2
1School of Biomedical Sciences, The Chinese University of Hong Kong, Shatin, New Territories, Hong Kong SAR, China.
Evaluating 10 tools for plotting precision-recall curves (PRC) and calculating area under the PRC (AUPRC) revealed significant differences. Some tools provide overly optimistic results, impacting classifier ranking in imbalanced datasets.
Area of Science:
- Bioinformatics
- Machine Learning
- Computational Biology
Background:
- Precision-recall curves (PRC) and area under the PRC (AUPRC) are crucial metrics for evaluating classification performance, especially in imbalanced datasets.
- These metrics are widely applied in fields like cancer diagnosis and cell type annotation.
- Over 3000 published studies have utilized tools for PRC plotting and AUPRC computation.
Purpose of the Study:
- To evaluate the performance and consistency of 10 popular tools used for generating precision-recall curves (PRC) and calculating the area under the PRC (AUPRC).
- To identify discrepancies in classifier ranking and potential biases in AUPRC values reported by different software tools.
Main Methods:
- A comparative analysis of 10 commonly used bioinformatics tools for PRC plotting and AUPRC calculation.
- Assessment of the consistency and accuracy of AUPRC values generated by these tools across various datasets.
- Evaluation of how different tools impact the ranking of classification models.
Main Results:
- Significant variability was observed in the area under the precision-recall curve (AUPRC) values computed by the evaluated tools.
- The ranking of classifiers differed substantially depending on the tool used for AUPRC calculation.
- Several tools were found to produce overly optimistic AUPRC results, potentially misrepresenting classifier performance.
Conclusions:
- The choice of tool for plotting precision-recall curves and computing area under the PRC significantly impacts classification performance evaluation.
- Users should exercise caution as some tools may yield inflated AUPRC values, leading to inaccurate assessments and rankings of classifiers.
- Further standardization or clear guidelines are needed for reliable AUPRC computation in machine learning applications, particularly for imbalanced data.
More Related Videos
16:23Automated, Quantitative Cognitive/Behavioral Screening of Mice: For Genetics, Pharmacology, Animal Cognition and Undergraduate Instruction
Published on: February 26, 2014
11:53Spatial Multiobjective Optimization of Agricultural Conservation Practices using a SWAT Model and an Evolutionary Algorithm
Published on: December 9, 2012
Related Concept Videos
Distribution Reliability and Automation
Propagation of Uncertainty from Systematic Error
Uncertainty: Overview
Introduction to R
Statgraphics
Propagation of Uncertainty from Random Error