Related Experiment Videos
Preventing Participant Leakage in Reissued Clinical AI Benchmarks: A Biomedical Informatics Evaluation Protocol and
Minoru Hattori1, Naoko Hasunuma1
1Hiroshima University, Japan, hiroshima.
Methods of Information in Medicine
|August 14, 2026
Summary
Participant leakage in clinical AI benchmarks inflates performance. Ensuring participant-disjoint evaluation is crucial for reliable benchmark reuse and accurate AI model assessment in biomedical informatics.
Area of Science:
- Biomedical informatics
- Artificial Intelligence (AI)
- Clinical AI benchmarks
Background:
- Clinical AI benchmarks are frequently reused, but successor releases are often treated as independent datasets.
- This practice can lead to benchmark-integrity threats if the same participants reappear across releases, causing evaluations to be internal while reported as external.
Purpose of the Study:
- Define successor-release participant leakage as a threat to benchmark integrity.
- Operationalize a pre-evaluation audit and leak-free protocol.
- Demonstrate the impact of leakage using the Distress Analysis Interview Corpus, Wizard-of-Oz condition (DAIC-WOZ) and Extended DAIC (E-DAIC) depression-screening benchmarks.
Main Methods:
- Audited release lineage and participant identity using persistent identifiers, content hashing, and label reconciliation.
- Compared leaky versus participant-disjoint protocols on held-out participants.
- Derived leak-free reference baselines across standard pipelines.
Main Results:
- All 189 DAIC-WOZ participants were found in E-DAIC with identical recordings and labels.
- A leaky acoustic protocol yielded an area under the receiver operating characteristic curve (AUROC) of 0.797, compared to 0.569 after deduplication.
- Leak-free reference baselines achieved an AUROC of approximately 0.60.
Conclusions:
- Successor-release participant leakage is a preventable evaluation failure in clinical AI benchmark reuse.
- Biomedical informatics studies must document release lineage, verify participant identity, reconcile labels, deduplicate across releases, and enforce participant-disjoint evaluation before pooling or comparing models.