Related Experiment Video
Updated: Apr 24, 2026

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
Published on: February 23, 2024
[Translated article] Reliability of artificial intelligence (ChatGPT) in the diagnosis and classification of tibial
C Castillejo1, M Zapatero1, J M Bogallo1
1Servicio de Cirugía Ortopédica y Traumatología, Hospital Universitario Costa del Sol, Universidad de Málaga, Marbella, Málaga, Spain.
Objective:
To compare the diagnostic and classification accuracy of tibial plateau fractures on simple radiographs among three groups: knee surgeons, resident physicians, and artificial intelligence (ChatGPT-4).
Methods:
An observational, descriptive, cross-sectional study with a control group was conducted on a prospective cohort of patients treated for tibial plateau fractures between 2020 and 2024. Anteroposterior radiographs were blindly evaluated by three groups - three knee surgeons, three resident physicians, and ChatGPT-4 - with fractures classified according to the Schatzker system. The reference standard was computed tomography (CT). The interobserver agreement was assessed using the Kappa statistic for fracture detection and the Ciccetti-weighted Kappa for fracture classification, with a 95% confidence interval. A significance level of p<0.01 was established.
Results:
A total of 387 radiographs were included, of which 129 showed tibial plateau fractures (classified according to Schatzker as follows: 7 type I, 28 type II, 5 type III, 16 type IV, 21 type V, and 52 type VI) and 258 were without fracture. The AI demonstrated the highest accuracy in fracture detection, achieving an absolute agreement of 99.5% and a Kappa of 0.98 (95% CI: 0.97-1.00, p<0.001), compared to 97% (κ=0.93, 95% CI: 0.91-0.95, p<0.001) for knee surgeons and 93% (κ=0.848, 95% CI: 0.81-0.88, p<0.001) for residents. In terms of interobserver variability for fracture diagnosis, the AI showed greater consistency than the human evaluators; however, for fracture classification, knee surgeons achieved a higher weighted Kappa (0.616, 95% CI: 0.554-0.679, p<0.001) compared to the AI (0.612, 95% CI: 0.502-0.722, p<0.001) and residents (0.572, 95% CI: 0.510-0.635, p<0.001).
Conclusions:
Artificial intelligence demonstrated notable accuracy in the detection of tibial plateau fractures, outperforming both residents and attending physicians in this specific task. However, in the classification of fractures using the Schatzker system, attending physicians achieved higher accuracy. These findings suggest that AI may serve as a valuable support tool in the diagnostic process, particularly in its early stages, complementing - but not replacing - the clinical judgment and experience of healthcare professionals.
Level Of Evidence:
Level III.
Diagnostic:
Cross-sectional descriptive study with control group.
