Related Experiment Video
Updated: Aug 6, 2026

11:21
Development of a Novel Internal Fixation Model for Rat Radial Fractures: Fracture Healing Assessment and Dorsal Root Ganglion Isolation
Published on: March 13, 2026
Can Artificial Intelligence Assess Distal Radius Fracture Stability?
Colby Newson1, Virginia Bailey1, Ava Chappell1,2
1Department of Plastic Surgery, University of Virginia Health System, Charlottesville, USA.
Summary
Large language models (LLMs) show limited accuracy in assessing distal radius fracture stability using the LaFontaine criteria. Current AI tools require significant validation before clinical use for fracture assessment.
Area of Science:
- Orthopedic surgery
- Artificial intelligence in medicine
- Radiographic interpretation
Background:
- Accurate distal radius fracture stability assessment is crucial for patient triage and specialist referral.
- The LaFontaine criteria aid in predicting fracture instability but are not routinely used.
- Multimodal large language models (LLMs) show promise for image interpretation, but their clinical utility is unproven.
Purpose of the Study:
- To evaluate the diagnostic accuracy and agreement of multimodal LLMs in assessing distal radius fracture stability.
- To compare LLM performance against hand surgeon consensus using the LaFontaine criteria.
Main Methods:
- A cross-sectional study assessed 20 distal radius fracture radiographs.
- Five hand surgeons established a reference standard for LaFontaine criteria and stability.
- Two LLMs (ChatGPT, Claude) were prompted to classify fracture stability and criteria.
Main Results:
- LLMs showed variable agreement with surgeon consensus on LaFontaine criteria.
- Both LLMs achieved similar overall accuracy (0.75) for fracture stability classification.
- ChatGPT demonstrated moderate agreement, while Claude showed fair agreement with clinicians.
Conclusions:
- Current multimodal LLMs have limited reliability for classifying distal radius fracture stability.
- LLM performance, particularly specificity, was insufficient for clinical decision-making.
- Further validation is necessary before LLMs can be used as adjunctive triage tools.
