Related Experiment Video
Updated: Jun 27, 2026

13:44
Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
Cross-Sensor and Cross-Population Generalization of Deep Learning Models for Digital Mammography: A Controlled
Somprasonk Gabbualoy1, Pattarapong Phasukkit1, Supan Tungjitkusolmun1
1School of Engineering, King Mongkut's Institute of Technology Ladkrabang, Bangkok 10520, Thailand.
Sensors (Basel, Switzerland)
|June 26, 2026
Summary
Deep learning models for digital mammography show decreased accuracy when transferred across different sensors and populations. Local validation and recalibration are crucial before clinical deployment of these AI tools.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Deep Learning for Mammography
- Computer-Aided Diagnosis
Background:
- Digital mammography AI models are increasingly used across diverse hospital settings.
- Variability in X-ray detector technology and patient demographics poses challenges for model generalizability.
- The performance of advanced mammography foundation models across different platforms remains under-tested.
Purpose of the Study:
- To evaluate the transferability and performance of state-of-the-art deep learning models for digital mammography.
- To assess whether models trained on one sensor and population maintain accuracy on others.
- To determine the necessity of local validation and recalibration for clinical deployment.
Main Methods:
- Fine-tuning five deep learning architectures (ResNet-50, DINOv2-B14, Rad-DINO, Mammo-CLIP B5, Mammo-FM) on the CBIS-DDSM dataset.
- Ablating a density-aware focal loss function to assess its impact on model performance.
- Evaluating model transfer to three external sensor cohorts (CMMD, DMID, MIAS) with statistical significance testing.
Main Results:
- ResNet-50 outperformed foundation models on the training dataset (CBIS-DDSM).
- Transferring models to external datasets resulted in significant AUC degradation (0.165-0.320 points).
- Sensitivity at 95% specificity decreased substantially upon transfer, highlighting performance collapse.
Conclusions:
- Mammography-specific pretraining did not guarantee superior in-distribution performance or robust transferability.
- No evaluated architecture achieved sufficient performance for defensible deployment without local validation.
- Sensor-stratified, population-stratified external validation and local recalibration are essential prerequisites for clinical use.