Related Experiment Video
Updated: Jun 13, 2026

Clinical Efficacy of an Innovative Multidimensional Traction Therapy in Moderate Adolescent Idiopathic Scoliosis
Published on: February 10, 2026
Clinical Evaluation of ChatGPT-5.3 Responses to Patient-Oriented Questions on Scoliosis: A Multidimensional Expert
1Department of Physical Medicine and Rehabilitation, Liv Hospital Vadistanbul, Ayazağa Mahallesi, Kemerburgaz Caddesi, Vadistanbul Park Etabı, 7F Blok, 34475 Istanbul, Turkey.
Abstract:
Background: The integration of large language models (LLMs) into healthcare has rapidly expanded, particularly in the domain of patient education. However, concerns remain regarding the accuracy, adequacy, and clinical safety of AI-generated medical information. This study aimed to systematically evaluate expert-perceived quality, content appropriateness, and potential clinical risk of responses generated by a GPT-5.3-based large language model to patient-oriented questions on scoliosis. Methods: A set of patient-oriented questions was developed based on common informational needs. Responses generated by the AI model were evaluated by a panel of experts using a multidimensional assessment framework, including general appropriateness, scientific accuracy, adequacy, and clarity. Content validity was assessed using the Content Validity Ratio (CVR) and Content Validity Index (CVI). In addition, clinical risk levels were categorized, and common issues in inappropriate responses were analyzed. CVR and CVI were used to quantify expert agreement regarding perceived content appropriateness, rather than to establish definitive factual correctness or guideline concordance. Results: All responses exceeded the predefined acceptable CVR threshold, and overall CVI values were high, indicating a high level of expert agreement regarding perceived response appropriateness. Clarity received the highest scores across all dimensions, whereas adequacy and scientific accuracy were relatively lower. Most responses were classified as harmless or low risk, and no high-risk responses were identified. The most frequently reported issue in inappropriate responses was insufficient information. Conclusions: Large language models may provide understandable and generally acceptable responses in a controlled expert-evaluation setting. However, these findings reflect expert-perceived appropriateness rather than definitive clinical validity or guideline concordance. AI-generated responses should therefore be used only as supportive educational tools under clinical oversight.
More Related Videos
07:01Clinical Efficacy of Ultrasound-Assisted Scoliosis-Specific Exercise in Mild-Grade Adolescent Idiopathic Scoliosis
Published on: December 2, 2025
06:28Biomechanical Changes Related to Low Back Pain: An Innovative Tool for Movement Pattern Assessment and Treatment Evaluation in Rehabilitation
Published on: December 13, 2024