Related Experiment Video
Updated: Jul 22, 2026

06:54
Clinical-oriented Three-dimensional Gait Analysis Method for Evaluating Gait Disorder
Published on: March 4, 2018
14.0K
Evaluating ChatGPT, Gemini and other Large Language Models (LLMs) in orthopaedic diagnostics: A prospective clinical
Stefano Pagano1, Luigi Strumolo2, Katrin Michalk1
1Department of Orthopaedic Surgery, University of Regensburg, Asklepios Klinikum, Bad Abbach, Germany.
Computational and Structural Biotechnology Journal
|January 24, 2025
Summary
Large Language Models (LLMs) show promise in diagnosing osteoarthritis. GPT-4o achieved high sensitivity using patient questionnaires, but medical oversight remains crucial for accurate diagnoses.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Diagnostics
- Natural Language Processing
Background:
- Large Language Models (LLMs) are emerging as potential tools in healthcare.
- This study investigates LLM diagnostic capabilities for hip or knee osteoarthritis (OA).
- Patient-reported data was used, excluding prior medical consultation.
Purpose of the Study:
- To evaluate the diagnostic sensitivity of various LLMs in detecting hip or knee osteoarthritis.
- To assess LLM performance using only patient-reported data from questionnaires.
- To compare LLM diagnoses against expert orthopaedic clinician diagnoses.
Main Methods:
- A prospective observational study involving 115 patients at an orthopaedic clinic.
- Patients completed a paper-based questionnaire on symptoms, history, and demographics.
- Five LLMs (including ChatGPT, Gemini, Llama, Gemma 2, Mistral-Nemo) were analyzed against clinician diagnoses.
Main Results:
- GPT-4o demonstrated the highest diagnostic sensitivity at 92.3%, outperforming other LLMs.
- Completeness of patient symptom reporting strongly predicted GPT-4o accuracy.
- Inter-model agreement varied, with GPT-4 versions showing moderate agreement and Llama-3.1 lower accuracy.
Conclusions:
- GPT-4o shows high accuracy and consistency in diagnosing OA from patient questionnaires.
- LLMs can serve as supplementary diagnostic tools, but require medical oversight.
- Further research is needed to enhance LLM diagnostic utility in healthcare.

