Related Experiment Video
Updated: Feb 4, 2026

Method and Instrumented Fixture for Femoral Fracture Testing in a Sideways Fall-on-the-Hip Position
Published on: August 17, 2017
ChatGPT-4o Mini Fabricates and Miscites Evidence for American Academy of Orthopaedic Surgeons Hip Fracture Clinical
David McCavitt1, Soroush Shabani1, Ashley Mulakaluri1
1Department of Orthopaedic Surgery, Keck School of Medicine of the University of Southern California, Los Angeles, California.
Background:
Generative artificial intelligence (AI) large language model (LLM) chatbots, such as ChatGPT, are increasingly used to answer medical questions. This study sought to assess the accuracy and quality of evidence cited in ChatGPT-4o mini responses to questions pertaining to hip fracture care.
Methods:
Prompt questions regarding hip fracture management that aligned with each of the 19 recommendations published in the American Academy of Orthopaedic Surgeons (AAOS) Clinical Practice Guideline (CPG) for Management of Hip Fractures in Older Adults were posed to the ChatGPT-4o mini LLM asynchronously by 4 independent medical student graders. Three prompt variations were applied for each recommendation, reflecting the perspectives of a physician, a patient, and a general information seeker. Graders then requested from the LLM a reference list with PubMed Identifier (PMID) numbers supporting each recommendation. Accuracy and clarity of responses were assessed using a standard rubric for overlap with CPG citations, fabrications, and inaccurate citations.
Results:
ChatGPT-4o mini returned 228 responses to prompts seeking advice on AAOS CPG hip management recommendations. 76.3% of responses were "accurate" to the CPG recommendation. 88.2% of responses received a clarity rating of "excellent". ChatGPT-4o mini provided 228 responses citing 2,556 publications when prompted for supporting evidence, of which 1.1% overlapped with AAOS CPG references, and 7.9% were fabricated. Of the publications cited by the LLM which exist in the PubMed index, 91.7% were given with incorrect authors, 91.5% incorrect titles, 91.4% incorrect pages, 91.0% incorrect PMIDs, 90.9% incorrect journals, 90.3% incorrect journal volumes, and 20.0% incorrect publication years. Responses for an AAOS CPG strong recommendation strength were significantly more likely to be "accurate" (p < 0.001), and responses for an AAOS CPG limited strength recommendation were significantly more likely to be "unsupported" (p < 0.001).
Conclusions:
ChatGPT-4o mini provided clear, moderately accurate responses with rampantly erroneous and occasionally fabricated citations to queries about hip fracture care derived from the AAOS Clinical Practice Guideline on Management of Hip Fractures in Older Adults.
Level Of Evidence:
Level V Therapeutic. See Instructions for Authors for a complete description of levels of evidence.
Related Concept Videos
The Evidence for Evolution
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Guidelines for Sketching a Curve
Guidelines for Nursing Documentation II
Timely documentation is crucial to ensure continuity of care for patients. Any delays in recording or reporting medical information can result in medical errors and even adverse patient outcomes. From medication administration to diagnostic test results, every detail must be accurately and promptly documented to provide the best possible care for patients.
Legal Guidelines for Documentation

