Related Experiment Video
Updated: Aug 20, 2025

13:44
Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
43.0K
External Validation of an Ensemble Model for Automated Mammography Interpretation by Artificial Intelligence
William Hsu1, Daniel S Hippe2, Noor Nakhaei1
1Medical and Imaging Informatics, Department of Radiological Sciences, David Geffen School of Medicine at University California, Los Angeles.
JAMA Network Open
|November 21, 2022
Summary
An ensemble artificial intelligence (AI) model for mammography screening showed lower performance in a diverse patient population compared to previous studies. This highlights the need for AI model transparency and fine-tuning for specific groups before clinical use.
Area of Science:
- Radiology and Medical Imaging
- Artificial Intelligence in Healthcare
- Machine Learning for Diagnostics
Background:
- Mammography screening programs face challenges due to a shortage of fellowship-trained breast radiologists.
- Artificial intelligence (AI) is being explored to enhance efficiency and diagnostic accuracy in mammography.
- External validation studies are crucial for assessing AI algorithm performance in diverse clinical settings.
Purpose of the Study:
- To externally validate an ensemble deep-learning model for breast cancer detection.
- To assess the model's performance using data from a high-volume, diverse patient population within an academic health system.
- To compare the AI model's performance against radiologist assessments.
Main Methods:
- An ensemble deep-learning model, combining outputs from 11 top AI models from the DREAM Mammography Challenge, was utilized.
- The study used retrospective screening mammography images from 37,317 examinations of women aged 40+ (2010-2020).
- Performance metrics including sensitivity, specificity, and AUROC were compared between the AI model, a combined AI-radiologist model (CEM+R), and radiologists alone.
Main Results:
- The ensemble deep-learning model achieved an AUROC of 0.85 in the UCLA cohort, lower than reported in other cohorts.
- The combined AI-radiologist model (CEM+R) showed similar sensitivity and specificity to radiologists overall.
- CEM+R demonstrated significantly lower sensitivity and specificity in women with a prior history of breast cancer and Hispanic women compared to radiologists.
Conclusions:
- The high performance of the ensemble deep-learning model did not generalize to a more diverse screening cohort, indicating potential underspecification.
- AI models require transparency and fine-tuning for specific target populations before widespread clinical adoption.
- Further research is needed to ensure equitable AI performance across diverse patient demographics in mammography screening.

