Related Experiment Video
Updated: Jan 22, 2026

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
Novel Artificial Intelligence Chest X-ray Diagnostics: A Quality Assessment of Their Agreement with Human Doctors in
Wolfram A Bosbach1,2, Luca Schoeni1, Jan Felix Senge3,4
1Bern University Hospital, University of Bern, Department of Nuclear Medicine, Inselspital, Switzerland, Bern.
Purpose:
The rising demand for radiology services calls for innovative solutions to sustain diagnostic quality and efficiency. This study evaluated the diagnostic agreement between two commercially available artificial intelligence (AI) chest X-ray systems and human radiologists during routine clinical practice.
Materials And Methods:
We retrospectively analyzed 279 chest X-rays (204 standing, 63 supine, 12 sitting) from a Swiss university hospital. Seven thoracic pathologies - cardiomegaly, consolidation, mediastinal mass, nodule, pleural effusion, pneumothorax, and pulmonary oedema - were assessed. Radiologists' routine reports were compared against Rayvolve (AZmed) and ChestView (Gleamer, both from Paris, France). A Python code, provided as open access supplement, calculated performance metrics, agreement measures, and effect size quantification.
Results:
Agreement between radiologists and AI ranged from moderate to almost perfect: Human-AZmed (Gwet's AC1: 0.47-0.72, moderate to substantial), and Human-Gleamer (Gwet's AC1: 0.56-0.96, moderate to almost perfect). Balanced accuracies ranged from 0.67-0.85 for Human-AZmed and 0.71-0.85 for Human-Gleamer, with peak performance for pleural effusion (0.85 both systems). Specificity consistently exceeded sensitivity across pathologies (0.70-0.98 vs 0.45-0.85). Common findings showed strong performance, pleural effusion (MCC 0.70-0.73), cardiomegaly (MCC 0.51), and consolidation (MCC 0.45-0.46). Rare pathologies demonstrated lower agreement, mediastinal mass, and nodules (MCC 0.23-0.31). Standing radiographs yielded superior agreement compared to supine studies. The two AI systems showed substantial inter-system agreement for consolidation and pleural effusion (balanced accuracy 0.81-0.84).
Conclusion:
Both commercial AI chest X-ray systems demonstrated comparable performance to human radiologists for common thoracic pathologies, with no meaningful differences between platforms. Performance was strongest for standing radiographs but declined for rare findings and supine studies. Position-dependent variability and reduced sensitivity for uncommon pathologies underscore the continued need for human oversight in clinical practice.
Key Points:
· AI systems matched radiologists for common chest X-ray findings.. · Standing radiographs achieved the highest diagnostic agreement.. · Rare pathologies showed weaker AI-human agreement.. · Supine studies reduced diagnostic performance.. · Human oversight remains essential in clinical practice..
Citation Format:
· Bosbach WA, Schoeni L, Senge JF et al. Novel Artificial Intelligence Chest X-ray Diagnostics: A Quality Assessment of Their Agreement with Human Doctors in Clinical Routine. Rofo 2025; DOI 10.1055/a-2778-3892.
Related Concept Videos
Myocarditis II: Clinical Features and Diagnostic Tests
Atherosclerosis II: Clinical Manifestations and Diagnostic Tests
Aneurysm II: Clinical Manifestations and Diagnostic Studies
Pericarditis II: Clinical Features and Diagnostic Tests
Mitral Stenosis II: Clinical features and Diagnostic Tests
Aortic Regurgitation II: Clinical Features and Diagnostic Tests

