Related Experiment Video
Updated: Jun 11, 2026

Clinical Efficacy of an Innovative Multidimensional Traction Therapy in Moderate Adolescent Idiopathic Scoliosis
Published on: February 10, 2026
An expert-led benchmark using common patient questions: evaluating large language models for adolescent idiopathic
Cemre Aydin1, Asli Beril Karakas2, Anil Murat Ozturk3
1Department of Orthopedics and Traumatology, Istanbul Bağcılar Education Research Hospital, 34200, Istanbul, Turkey.
Aim/Background:
Large language models (LLMs) are increasingly used by patients to obtain medical information. Adolescent idiopathic scoliosis (AIS), a chronic condition requiring long-term monitoring and treatment decisions, generates substantial demand for reliable and understandable patient education. Although LLMs may function as accessible explanatory tools, their suitability for patient-oriented use remains uncertain. This study aimed to perform an expert-led, patient-centered evaluation of two widely accessible LLMs, Claude Sonnet 4.5 and GPT 5.2, focusing on their ability to deliver accurate, clear, and conceptually adequate responses to common AIS-related patient questions.
Methods:
A cross-sectional comparative design was used with 100 high-frequency patient questions covering ten clinical domains. Responses generated by both models using standardized zero-shot prompts were independently assessed by expert clinicians: factual accuracy by three raters (two orthopedic spine surgeons and one senior pediatric physiotherapist), and clarity and conceptual coverage by two raters (one surgeon and the physiotherapist). A structured evaluation framework examined three dichotomous dimensions relevant to patient education: factual accuracy, clarity and understandability, and conceptual coverage. Model performances were compared using McNemar's test, and inter-model agreement was assessed with Krippendorff's alpha.
Results:
Both models demonstrated equally high factual accuracy (91%). However, clarity was limited, with only one-third of responses rated as sufficiently understandable. A significant difference was observed in conceptual coverage, with Claude Sonnet 4.5 outperforming GPT 5.2 (46% vs. 29%, p = 0.012), particularly in domains requiring integrative explanations.
Conclusion:
Despite strong factual accuracy, current LLMs show deficiencies in clarity and conceptual depth, limiting their reliability as standalone patient education tools for AIS. These findings highlight the necessity of clinician mediation and the importance of patient-centered evaluation criteria before clinical adoption.
Clinical Trial Registration:
As this study is not a clinical trial, clinical trial registration is not applicable.
More Related Videos
07:01Clinical Efficacy of Ultrasound-Assisted Scoliosis-Specific Exercise in Mild-Grade Adolescent Idiopathic Scoliosis
Published on: December 2, 2025
08:08Cell-based Assay Protocol for the Prognostic Prediction of Idiopathic Scoliosis Using Cellular Dielectric Spectroscopy
Published on: October 16, 2013