Related Experiment Video
Updated: Aug 6, 2026

Software-Assisted Quantitative Measurement of Osteoarthritic Subchondral Bone Thickness
Published on: March 18, 2022
Performance of multimodal large language models versus clinicians for radiographic knee osteoarthritis grading: A
Asli Irmak Akdogan1, Efe Kemal Akdogan2, Mehmet Fatih Tumer3
1Department of Radiology, Izmir Katip Celebi University, Ataturk Training and Research Hospital, Izmir, Turkey. irmakbiranci@gmail.com.
Objective:
To compare the performance of clinicians and two generations of multimodal large language models (LLMs) in Kellgren-Lawrence (KL) grading of knee osteoarthritis (KOA), including feature-level assessment and intraobserver repeatability.
Materials And Methods:
In this retrospective single-center study, 348 knee radiographs were graded by a senior musculoskeletal radiologist (reference standard), a radiologist, an orthopedic surgeon, a radiology resident, and LLMs (ChatGPT-4o and ChatGPT-5.0). Binary KOA detection (KL 0-1 vs ≥ 2), feature-level interpretation (joint space narrowing, osteophytes, subchondral sclerosis), and intraobserver repeatability were evaluated. Agreement metrics included weighted κ, accuracy, and standard diagnostic measures.
Results:
Agreement with the reference standard was highest for the radiologist (κ = 0.87), followed by the orthopedic surgeon and radiology resident. Both LLMs demonstrated moderate agreement, with ChatGPT-5.0 outperforming ChatGPT-4o. For binary KOA detection, ChatGPT-5.0 showed very high sensitivity (0.96) but reduced specificity. Per-grade classification was most accurate for KL 0 and KL 4, but remained limited for KL 1-2. Feature-level concordance was modest across all radiographic findings. Intraobserver repeatability was highest for the reference reader (κ = 0.881), followed by the orthopedic surgeon (κ = 0.634) and radiologist (κ = 0.626), while lower agreement was observed for the resident (κ = 0.478) and LLMs, with ChatGPT-5.0 showing higher consistency than ChatGPT-4o (κ = 0.591 vs. 0.485).
Conclusion:
Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility. Current multimodal LLMs show high sensitivity but limited specificity and are not suitable for standalone radiographic KOA assessment.