Related Experiment Video
Updated: Sep 27, 2026

Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
Ten Years of Artificial Intelligence in Screening Mammography: A Systematic Review and Meta-Analysis of Diagnostic
Sebastian Ciurescu1, Victor Buciu1, Diana-Gabriela Ilaș2
1Doctoral School in Medicine, Victor Babeș University of Medicine and Pharmacy, 300041 Timisoara, Romania.
Abstract:
Background: Deep-learning artificial intelligence (AI) for mammographic screening moved from proof of concept to randomised evaluation in a single decade. We reviewed and meta-analysed its diagnostic accuracy and its effect on screening programmes, covering the literature published between 2015 and 2025. Methods: PubMed and Europe PMC were searched from 1 January 2015 to 31 December 2025, supplemented by ClinicalTrials.gov and by forward and backward citation searching (PRISMA 2020, PRISMA-DTA, PRISMA-S). Eligible studies evaluated a deep-learning system for cancer detection or triage in a screening population against a histopathological reference standard. Four syntheses were performed: standalone accuracy pooled on the logit-AUC scale (A), a bivariate sensitivity-specificity model (A2), cancer detection rate ratio for AI-integrated versus standard reading (B), and recall rate ratio (C). Random-effects models used restricted maximum likelihood with Knapp-Hartung intervals. Risk of bias was assessed with QUADAS-2 and QUADAS-C, and certainty with GRADE. Results: Twenty-six studies (27 reports, 2019-2025) were included. Pooled standalone AUC across 14 studies and 1,214,885 examinations was 0.890 (95% CI 0.858-0.915), with I2 = 97.4% and a 95% prediction interval of 0.731-0.960. Neither publication year (p = 0.79) nor enriched versus consecutive sampling (p = 0.94) explained this dispersion in meta-regression. The bivariate model (k = 9) gave a summary sensitivity of 73.3% (64.5-80.6) at a specificity of 92.4% (86.8-95.8). Across randomised and paired prospective trials (k = 3), the pooled detection rate ratio was 1.13 (0.83-1.55), with the MASAI randomised trial alone reporting 1.29 (1.09-1.51) and a 44% reduction in screen reading. Non-randomised implementation studies (k = 5), which include a 463,094-women German programme evaluation, pooled to 1.22 (1.08-1.37) with little dispersion (I2 = 19.0%). Recall changed little overall (0.95, 0.81-1.12). Certainty was very low for accuracy outcomes and moderate for the single randomised trial. Conclusions: A decade of evidence supports AI as a second reader and triage tool in organised screening, not as an autonomous replacement for the radiologist. Pooled accuracy is high on average but so dispersed that it cannot be transferred to a new programme; local validation before deployment remains necessary. The larger and more precise detection gains come from non-randomised designs, which is the pattern confounding would produce, so the randomised evidence remains the anchor. Interval-cancer and mortality endpoints are still awaited.
