Related Experiment Video
Updated: Dec 9, 2025

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.9K
An efficient variance estimator of AUC and its applications to binary classification.
1Department of Mathematics, Wellesley College, Wellesley, Massachusetts.
Statistics in Medicine
|September 11, 2020
Summary
We developed a new, unbiased variance estimator for the area under the ROC curve (AUC) to accurately assess classifier performance. This efficient method is significantly faster than existing techniques.
Area of Science:
- Statistics
- Machine Learning
- Medical Data Analysis
Background:
- Area under the ROC curve (AUC) is a key metric for binary classifier performance.
- Sampling variation necessitates assessing the reliability of AUC estimates.
- Current methods for AUC variance estimation can be computationally intensive.
Purpose of the Study:
- To develop an unbiased variance estimator for AUC estimates.
- To generalize the estimator for K-sample U-statistics.
- To provide a computationally efficient method for AUC variance estimation.
Main Methods:
- Extended Wang and Lindsay's proposal to a two-sample U-statistic form.
- Employed a partition-resampling scheme for computational efficiency.
- Validated the estimator through simulation studies and real-world medical datasets.
Main Results:
- The proposed AUC variance estimator is unbiased.
- It demonstrates comparable or superior performance to jackknife and bootstrap methods.
- Achieved computational times 10-30 times faster than existing methods.
Conclusions:
- The developed AUC variance estimator offers an efficient and accurate alternative.
- It is applicable for model selection (e.g., one-standard-error rule) and confidence interval construction.
- The method shows practical utility in medical sciences.
Related Concept Videos
Receiver Operating Characteristic Plot
418
A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
418
Variance
11.6K
The deviations show how spread out the data are about the mean. A positive deviation occurs when the data value exceeds the mean, whereas a negative deviation occurs when the data value is less than the mean. If the deviations are added, the sum is always zero. So one cannot simply add the deviations to get the data spread. By squaring the deviations, the numbers are made positive; thus, their sum will also be positive.
The standard deviation measures the spread in the same units as the data....
The standard deviation measures the spread in the same units as the data....
11.6K
Bioequivalence Data: Statistical Interpretation
110
Body:The statistical interpretation of bioequivalence data is a significant aspect of pharmaceutical research. Bioequivalence refers to the absence of any significant difference in the rate and extent to which the active ingredient in pharmaceutical products becomes available at the site of drug action when administered at the same molar dose under similar conditions. This helps determine if different drug products have similar absorption rates, ensuring their interchangeability.Statistical...
110
One-Way ANOVA: Equal Sample Sizes
3.9K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.9K
Kaplan-Meier Approach
450
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...
450
What is an ANOVA?
8.9K
The Analysis of Variance or ANOVA is a statistical test developed by Ronald Fisher in 1918. It is performed on three or more samples to check for equality between their means.
Before performing ANOVA, one must ensure that the samples used for this analysis have three crucial characteristics or statistical assumptions. The first assumption states that the samples should be drawn from normally distributed samples, while the second requires that all the drawn samples should be randomly and...
Before performing ANOVA, one must ensure that the samples used for this analysis have three crucial characteristics or statistical assumptions. The first assumption states that the samples should be drawn from normally distributed samples, while the second requires that all the drawn samples should be randomly and...
8.9K

