Related Experiment Video
Updated: Dec 26, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Monte Carlo studies of bootstrap variability in ROC analysis with data dependency.
Jin Chu Wu1, Alvin F Martin1, Raghu N Kacker1
1National Institute of Standards and Technology, Gaithersburg, Maryland 20899, USA.
This study addresses data dependency in Receiver Operating Characteristic (ROC) analysis using a two-layer bootstrap method. It recommends 2,000 bootstrap replications for accurate statistical analysis when dealing with limited, dependent data.
Area of Science:
- Statistics
- Biostatistics
- Machine Learning
Background:
- Receiver Operating Characteristic (ROC) analysis is crucial for classifier performance evaluation across disciplines.
- Data dependency, arising from repeated subject use, is common due to resource limitations.
- Estimating standard errors for ROC statistics with dependent data requires specialized methods.
Purpose of the Study:
- To determine the optimal number of bootstrap replications for ROC analysis with data dependency.
- To reduce bootstrap variance and enhance computational accuracy.
- To provide guidance for robust statistical inference in ROC analysis.
Main Methods:
- A two-layer data structure was employed to handle data dependency.
- Nonparametric two-sample two-layer bootstrap was utilized for standard error estimation.
- Monte Carlo studies assessed bootstrap variability to find appropriate replication numbers.
Main Results:
- The study investigated bootstrap variability in ROC analysis with dependent data.
- A specific number of bootstrap replications was identified as appropriate.
- A tolerance of 0.02 for the coefficient of variation was used.
Conclusions:
- 2,000 bootstrap replications are suggested for ROC analysis with data dependency.
- This recommendation ensures accuracy and reduces variance in statistical computations.
- The findings support reliable classifier evaluation in resource-constrained scenarios.
More Related Videos
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
04:35Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Related Concept Videos
Bootstrapping
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Friedman Two-way Analysis of Variance by Ranks
Contingency Table
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...