Leveraging Large Language Models for Accurate AO Fracture Classification from CT Text Reports

Markus Mergen1, Daniel Spitzl2, Conrad Ketzer3

  • 1Department of Diagnostic and Interventional Radiology, School of Medicine, TUM University Hospital, Technical University of Munich, 81675, Munich, Germany.

Summary

Large language models (LLMs) show potential for fracture classification in radiology reports. ChatGPT-4o and AmbossGPT performed best, but all LLMs need further refinement for detailed subtype accuracy.