Related Experiment Video
Updated: Sep 9, 2026

Fine-Tuning Large Language Models Using Entity Hallucination Index for Text Summarization
Published on: January 9, 2026
Select large language models outperform hip preservation experts on consensus-based hip preservation questionnaire
Karter Morris1, Felix Öttl2, James A Pruneski3
1School of Medicine, Texas Tech University Health Sciences Center, Lubbock, Texas, USA.
Purpose:
Artificial intelligence (AI) is increasingly utilized in medical education and clinical contexts, yet few studies compare the performance of large language models (LLMs) to subspecialized experts in providing guideline-based medical information on hip preservation. The purpose of this study was to evaluate the performance of three LLMs compared to a panel of international hip preservation experts in answering guideline-based questions related to femoroacetabular impingement syndrome, hip dysplasia and microinstability of the hip.
Methods:
A 21-item questionnaire was developed based on published consensus guidelines. The survey was distributed to a panel of hip preservation specialists identified through the professional network of the senior author. Ten experts responded and were included in the analysis. Three LLMs (ChatGPT 5.2, Gemini 3 and Claude 4.5 Sonnet) were prompted using the same questionnaire, with each LLM performing three runs per item. Outcomes included overall accuracy, percent agreement, Fleiss' κ, generalized linear mixed-effects modelling and qualitative assessment of AI answer justifications.
Results:
Expert accuracy was 90.5% (95% confidence interval [CI] 87.1-93.9), compared to 100% (p = 0.004) for Gemini, 98.4% (p = 0.016) for ChatGPT and 96.8% (p = 0.053) for Claude. Expert percent agreement was 42.9% and Fleiss' κ was 0.769; alternatively, AI intra-item percent agreement was 100% (Gemini) and 95.2% (ChatGPT and Claude). ChatGPT and Claude provided thorough justifications for even incorrect responses, and Gemini demonstrated formatting deviations despite 100% accuracy.
Conclusion:
The three LLMs demonstrated high accuracy and consistency when answering the hip preservation questionnaire, with two of the LLMs statistically outperforming the expert panel. In structured, verifiable question sets, the ability of newer LLMs to accurately and consistently respond to consensus-based questions is improving compared to prior reports. LLMs are likely to serve as an adjunct in orthopaedic education and practice, and limitations to AI's implementation into practice should be continuously and rigorously explored.
Level Of Evidence:
Level V.
Related Concept Videos
Methods of Documentation II: POMR
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities