Related Experiment Video
Updated: Sep 20, 2026

Autologous Microfractured and Purified Adipose Tissue for Arthroscopic Management of Osteochondral Lesions of the Talus
Published on: January 23, 2018
Large Language Models for Ankle Fracture Classification and Management Prediction from Routine Clinical
Benjamin Schwarberg1, Conrad Ketzer2, Bosse Thiel2
1Department of Diagnostic and Interventional Radiology, TUM School of Medicine and Health, TUM University Hospital, Technical University of Munich, Rechts Der Isar, Munich, Germany. Benjamin-Schwarberg@gmx.de.
Abstract:
Ankle fractures are among the most common injuries in trauma surgery and require accurate classification, consistent documentation, and individualized management. Large language models (LLMs) offer the potential to translate unstructured clinical text reports into structured, management-related information, yet their role in orthopedic trauma workflows remains insufficiently defined. In this retrospective study, the performance of Llama 3.1 (70B Instruct) was evaluated using routine radiology reports and clinical documentation from 54 patients with acute ankle fractures. Outputs were compared with reference standards derived from routine clinical documentation for Weber fracture classification, operative versus nonoperative management, and surgical procedure category. Four prompting strategies were systematically assessed. Accuracy for Weber classification ranged from 0.759 to 0.815, well above majority-class baseline (0.556), with strong performance for Weber B fractures and most errors occurring between adjacent categories. Macro-averaged F1 across the three Weber classes ranged from 0.744 to 0.822. Operative management prediction reached accuracies of 0.833 to 0.963 (sensitivity 0.872 to 1.000, specificity 0.571 to 0.714). Procedure category prediction demonstrated lower accuracy (0.426 to 0.610 depending on the endpoint), largely at or below a majority-class baseline (0.553), reflecting the limitations of text-only input for operative planning. These findings suggest that LLMs can extract structured information from routine clinical documentation, performing well for standardized classification but less reliably for procedure-level prediction. Prospective and multimodal validation is needed before clinical implementation.