Related Experiment Video
Updated: Aug 29, 2026

Protocol for the Evaluation of MRI Artifacts Caused by Metal Implants to Assess the Suitability of Implants and the Vulnerability of Pulse Sequences
Published on: May 17, 2018
Evaluation metrics for synthetic medical imaging
Daniel de Wilde1, Benjamin Schärli1, Kym Ackermann1
1Machine Intelligence in Clinical Neuroscience & Microsurgical Neuroanatomy (MICN) Laboratory, Department of Neurosurgery, Clinical Neuroscience Center, University Hospital Zurich, University of Zurich, Zurich, Switzerland.
Background:
Advances in generative artificial intelligence (AI) have accelerated the development and application of synthetic medical imaging. Despite this rapid progress, the evaluation of synthetic medical images remains heterogeneous, with numerous metrics proposed to assess fidelity, realism, diversity, and clinical validity. Currently, no standardized framework exists to guide the selection, interpretation, or comparison of these metrics, limiting reproducibility and cross-study comparability. This systematic review aims to comprehensively summarize and categorize existing metrics used to assess these complementary dimensions of synthetic medical images.
Methods:
A systematic review was conducted in accordance with PRISMA guidelines. PubMed/MEDLINE, EMBASE, Scopus, and arXiv were searched for studies published between 2015 and April 30, 2025, supplemented by citation screening of included studies. Eligible studies were full-text articles that applied or proposed metrics to evaluate the fidelity, realism, diversity, and/or clinical validity in synthetic medical images.
Results:
A total of 47 studies were included. Evaluation practices were highly heterogeneous. Expert evaluation (n = 25, 53%) and reference-based evaluations were most common (n = 25, 53%), followed by no-reference metrics (n = 24, 51%), and task-based evaluations (n = 24, 51%). The most commonly used individual metrics were peak signal-to-noise ratio (PSNR) (n = 16, 34%), structural similarity index (SSIM) (n = 15, 32%), mean absolute error (MAE) (n = 12, 26%), and Fréchet Inception Distance (FID) (n = 12, 26%).
Conclusion:
Evaluation strategies for synthetic medical imaging showed substantial variability and no single metric captured fidelity, realism, diversity, and clinical validity simultaneously. Metric choice is often dictated by data availability rather than clinical purpose. A task-specific, layered evaluation framework could improve comparability and facilitate clinical adoption.