Consensus-Level and Cluster-Adjusted Evaluation of a Large Language Model for Diagnostic Extraction from

Wolfram A Bosbach1, Elham Montazeri1, Jan F Senge2,3

  • 1Department of Nuclear Medicine, Inselspital, Bern University Hospital, University of Bern, 3010 Bern, Switzerland.

Summary

Large language models like ChatGPT-4.0 show high accuracy in extracting diagnoses from musculoskeletal radiology reports, comparable to experienced radiologists. Individual AI performance slightly exceeds human readers, but consensus-level accuracy is similar for both.

Related Concept Videos