Related Experiment Video
Updated: Sep 14, 2026

Treatment of Ankle Osteoarthritis with Total Ankle Replacement Through a Lateral Transfibular Approach
Published on: January 24, 2018
High Accuracy, Questionable References: Large Language Models' Performance and Citation Reliability in Foot and Ankle
Benjamin Nieves-Lopez1, Andrea Fabregas2, Gonzalo F Del Rio Montesinos2
1Department of Orthopaedic Surgery, University of Puerto Rico, Medical Sciences Campus, San Juan, Puerto Rico.
Background:
Artificial intelligence large language models (LLMs) are increasingly used for medical education and decision support in orthopaedics, yet their performance on foot and ankle surgery board-style questions remains underexplored. Additionally, the capability of LLMs to provide reliable references and source material for education has not been evaluated. Therefore, this study aimed to compare the accuracy and citation reliability of ChatGPT‑5.4, Gemini‑3, and Copilot on foot and ankle surgery Self-Assessment Examination (SAE) questions.
Methods:
A total of 192 multiple-choice questions (64 text-only, 128 image-based) from OrthoBullets Foot and Ankle SAE forms D and E were entered into ChatGPT‑5.4, Gemini‑3, and Copilot; video-based items were excluded because of Copilot limitations. Accuracy was compared overall and by question type and imaging modality. References volunteered or prompted were classified as real, partially fabricated, or completely fabricated and compared across models and by answer correctness. Statistical significance was set at a P value <.05.
Results:
ChatGPT‑5.4 achieved the highest overall accuracy (89.6%) vs Gemini‑3 (76.6%) and Copilot (74.0%). ChatGPT‑5.4 and Gemini‑3 showed accuracy that did not differ significantly on text-only vs image-based questions, whereas Copilot performed significantly worse on image-based items (P = .028). With the numbers available, no significant difference in accuracy could be detected by imaging modality for any model. ChatGPT‑5.4, Gemini‑3, and Copilot fabricated 22.9%, 16.1%, and 2.8% of all references, respectively; most ChatGPT‑5.4 and Gemini‑3 fabrications were partially fabricated, whereas all Copilot fabrications were completely fabricated. ChatGPT‑5.4 fabricated significantly more references when it answered incorrectly (P = .013).
Conclusion:
ChatGPT‑5.4 showed the highest accuracy on foot and ankle SAE questions, with diminished image-related performance gaps relative to prior generations, but also the highest rate of reference fabrication. All models produced fabricated citations, underscoring the need to verify AI-generated references and to use LLMs as adjunct rather than primary educational tools in foot and ankle surgery.

