Related Experiment Video
Updated: May 5, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Prioritizing cases from a multi-institutional cohort for a dataset of pathologist annotations
Victor Garcia1, Emma Gardecki1, Stephanie Jou2
1U.S. Food and Drug Administration, Center for Devices and Radiological Health, Office of Science and Engineering Laboratories, Division of Imaging, Diagnostics, and Software Reliability, Silver Spring, MD, United States of America.
We developed a method to create a validation dataset for artificial intelligence models assessing tumor-infiltrating lymphocytes in breast cancer. This approach prioritizes underrepresented patient groups for more equitable AI development.
Area of Science:
- Oncology
- Computational Pathology
- Artificial Intelligence
Background:
- Standardized validation datasets are crucial for comparing artificial intelligence and machine learning (AI/ML) models in cancer research.
- Stromal tumor-infiltrating lymphocytes (sTILs) are important prognostic markers in triple-negative breast cancer (TNBC).
- Developing robust AI/ML models for sTILs assessment requires high-quality, diverse validation data.
Purpose of the Study:
- To create a comprehensive validation dataset for AI/ML models assessing sTILs in TNBC.
- To implement a novel case prioritization method to ensure representation of diverse patient subgroups.
- To facilitate direct comparison of AI/ML model performance across different research groups.
Main Methods:
- Obtained whole slide images (WSIs) and clinical metadata for TNBC core biopsies from two academic medical centers.
- Selected regions of interest (ROIs) targeting diverse tissue morphologies and sTILs densities.
- Implemented a hierarchical rank-sort method for case prioritization, focusing on underrepresented clinical factors.
Main Results:
- Compiled data from 122 glass slides across 105 unique TNBC patients.
- Improved the skewness of sTILs density distribution from 0.60 to 0.46 through case prioritization.
- Increased the entropy of sTILs density bins from 1.20 to 1.24, enhancing data diversity.
Conclusions:
- The developed prioritization method effectively enhances the representation of underrepresented patient subgroups in the validation dataset.
- This approach is vital for creating a robust and equitable dataset for training and validating AI/ML models in TNBC research.
- The methodology described facilitates the creation of pivotal studies for AI/ML model development in digital pathology.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Steps in Outbreak Investigation
Comparing the Survival Analysis of Two or More Groups
Statistical Software for Data Analysis and Clinical Trials
Hazard Ratio
For example, in a clinical trial...
Investigation of Disease Outbreaks

