Related Experiment Video
Updated: Oct 9, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Nonparametric inference of the area under ROC curve under two-phase cluster sampling.
1Department of Epidemiology and Biostatistics, College of Public Health, University of South Florida, Tampa, Florida, USA.
This study introduces a new method for estimating the area under the ROC curve (AUC) that accounts for both verification bias and clustering in medical studies. The proposed approach provides a more accurate variance estimate, crucial for reliable diagnostic test evaluation.
Area of Science:
- Biostatistics
- Medical Informatics
- Diagnostic Test Evaluation
Background:
- Nonparametric inference for the area under the ROC curve (AUC) is established for verification bias or clustering individually.
- Existing methods fail when both verification bias and clustering are present, common in two-phase studies with multiple test results per subject.
- Accurate AUC inference requires accounting for both verification bias and intra-cluster correlation.
Purpose of the Study:
- To develop a nonparametric method for AUC inference that simultaneously addresses verification bias and clustering.
- To propose an inverse probability weighting (IPW) AUC estimator robust to these combined biases.
- To derive a variance formula that correctly accounts for intra-cluster correlations.
Main Methods:
- Proposed an Inverse Probability Weighting (IPW) AUC estimator.
- Developed a variance formula incorporating intra-cluster correlations for disease status and test results.
- Utilized a simulation study to evaluate the proposed method's performance.
Main Results:
- Methods assuming independence between subjects and test results underestimate the true variance of the IPW AUC estimator when intra-cluster correlations exist.
- The proposed method provides a consistent variance estimate for the IPW AUC estimator.
- The simulation results demonstrate the superiority of the proposed approach in handling combined biases.
Conclusions:
- The proposed IPW AUC estimator and variance formula effectively address the challenges of verification bias and clustering in diagnostic accuracy studies.
- Accounting for intra-cluster correlations is essential for accurate variance estimation in such complex study designs.
- This method enhances the reliability of AUC inference in real-world clinical and epidemiological research.
More Related Videos
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
Related Concept Videos
Receiver Operating Characteristic Plot
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Comparing the Survival Analysis of Two or More Groups
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...