Leveraging Large Language Models for Accurate AO Fracture Classification from CT Text Reports
Markus Mergen1, Daniel Spitzl2, Conrad Ketzer3
1Department of Diagnostic and Interventional Radiology, School of Medicine, TUM University Hospital, Technical University of Munich, 81675, Munich, Germany.
Journal of Imaging Informatics in Medicine
|July 7, 2025
Summary
Large language models (LLMs) show potential for fracture classification in radiology reports. ChatGPT-4o and AmbossGPT performed best, but all LLMs need further refinement for detailed subtype accuracy.
Area of Science:
- Radiology
- Artificial Intelligence
- Medical Informatics
Background:
- Large language models (LLMs) offer potential for analyzing complex radiological reports.
- LLMs can aid clinicians by integrating diagnostic criteria into classifications.
- Rigorous validation is crucial for clinical adoption of LLMs in advanced radiology.
Purpose of the Study:
- To evaluate the performance of four LLMs (ChatGPT-4o, AmbossGPT, Claude 3.5 Sonnet, Gemini 2.0 Flash) in classifying fractures using the AO classification system.
- To compare the accuracy of these LLMs on a dataset of CT reports.
- To identify LLM strengths and limitations in fracture classification.
Main Methods:
- Retrospective analysis of 292 fictitious CT reports with 310 fractures.
- LLMs were tasked with AO fracture classification based on report text.
- Performance was measured against ground truth labels, with accuracy analyzed by fracture type and subtype.
Main Results:
- ChatGPT-4o (74.6%) and AmbossGPT (74.3%) showed the highest overall accuracy.
- Claude 3.5 Sonnet (69.5%) and Gemini 2.0 Flash (62.7%) had lower accuracy.
- Bone recognition was high (90-99%), but fracture subtype classification accuracy was limited (71-77%).
Conclusions:
- LLMs demonstrate potential for assisting radiologists with initial fracture classification, especially in high-volume settings.
- Current LLM performance is inconsistent for detailed fracture subtype classification.
- Further refinement and validation are necessary before widespread clinical integration for advanced diagnostic workflows.


