Related Experiment Video
Updated: Jun 13, 2026

12:50
Lesion Explorer: A Video-guided, Standardized Protocol for Accurate and Reliable MRI-derived Volumetrics in Alzheimer's Disease and Normal Elderly
Published on: April 14, 2014
Consensus-Level and Cluster-Adjusted Evaluation of a Large Language Model for Diagnostic Extraction from
Wolfram A Bosbach1, Elham Montazeri1, Jan F Senge2,3
1Department of Nuclear Medicine, Inselspital, Bern University Hospital, University of Bern, 3010 Bern, Switzerland.
Diagnostics (Basel, Switzerland)
|June 12, 2026
Summary
Large language models like ChatGPT-4.0 show high accuracy in extracting diagnoses from musculoskeletal radiology reports, comparable to experienced radiologists. Individual AI performance slightly exceeds human readers, but consensus-level accuracy is similar for both.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Radiology Report Analysis
- Natural Language Processing in Healthcare
Background:
- Administrative workload in radiology can be reduced by automating diagnostic extraction from text reports using large language models (LLMs).
- Evaluating the accuracy of LLMs in clinical settings is crucial for their adoption.
- Musculoskeletal (MSK) radiology reports contain complex diagnostic information requiring precise extraction.
Purpose of the Study:
- To evaluate the diagnostic accuracy of ChatGPT-4.0 in extracting diagnoses from MSK radiology text reports.
- To compare the performance of ChatGPT-4.0 against experienced human readers.
- To analyze performance using cluster-adjusted and consensus-level statistical methods.
Main Methods:
- Analysis of 23 multimodal MSK imaging cases (X-ray, ultrasound, CT, MRI).
- Ten human readers and ChatGPT-4.0 (10 iterations) provided primary and secondary diagnoses from six predefined options.
- Individual-reader analysis using cluster-adjusted generalized estimating equations (GEE) and case-level analysis using majority consensus with exact McNemar testing.
Main Results:
- ChatGPT-4.0 achieved higher accuracy (0.957) for primary diagnoses compared to human readers (0.865).
- At the consensus level, discordance was low (8.7%), with no significant difference between AI and human methods.
- Combined primary and secondary diagnoses showed perfect consensus (23/23) for both AI and humans, with almost perfect interrater reliability (Gwet's AC1 = 0.836-0.927).
Conclusions:
- ChatGPT-4.0 demonstrates diagnostic accuracy comparable to experienced radiologists in structured MSK text reports.
- Individual AI reader advantage diminishes at the consensus level, suggesting AI can augment, not replace, human expertise.
- Variability is primarily case-driven, indicating future validation studies should prioritize case numbers over reader numbers for robust findings.
