Related Experiment Video
Updated: Jan 9, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Assessing the generalizability of artificial intelligence in radiology: a systematic review of performance across
Muhammad Umer Suleman1, Muhammad Mursaleen1, Umer Khalil1
1Department of Internal Medicine, Ayub Medical College, Abbottabad, Pakistan.
Introduction:
Artificial intelligence (AI) applications in diagnostic radiology have demonstrated remarkable accuracy on institutional datasets. However, concerns about external generalizability performance when models encounter data from different hospitals, scanners, or patient populations remain a major barrier to clinical deployment.
Methods:
We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses and Assessing the Methodological Quality of Systematic Reviews -compliant systematic review of peer-reviewed studies (January 2022 - June 2025) reporting both internal and external validation of AI diagnostic models applied to computed tomography, magnetic resonance imaging, or X-rays. The review was registered in the PROSPERO database. PubMed and Embase searches identified 342 records; after de-duplication, screening, and eligibility assessment, six studies met our inclusion criteria. These studies addressed diverse diagnostic tasks using deep learning architectures (3D Convolutional Neural Networks, Generative Adversarial Networks-augmented models, no-new-U-Net ensembles, and regulatory-cleared systems).
Results:
Internal-validation area under the curve (AUC) ranged from 0.76 to 0.95; sensitivities were generally >85% and specificities >68%. In external validation, performance declined modestly in AUC (median drop ~0.03), with larger decreases in specificity (up to ~24 percentage points). Quality Assessment of Diagnostic Accuracy Studies - Version 2 assessment revealed low overall risk of bias in five studies; one study had high patient-selection bias, and another had unclear sampling. Methods that appeared to enhance generalizability included multicenter training, data augmentation with generative adversarial networks, and incorporation of clinical variables.
Conclusion:
AI models in radiology tend to underperform on external data despite strong internal performance. Mandatory external validation on diverse cohorts and cautious clinical integration are recommended.
More Related Videos
05:49Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
Published on: February 23, 2024
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Magnetic Resonance Imaging
Radiological Investigation I: X-ray and CT
Imaging Studies IV: Magnetic Resonance Imaging