Related Experiment Video
Updated: Apr 6, 2026

Integrating Augmented Reality Tools in Breast Cancer Related Lymphedema Prognostication and Diagnosis
Published on: February 6, 2020
Performance of ChatGPT-4o, Gemini 2.0 Pro, and DeepSeek-V3 in Patient-Facing Information on Chest Wall Deformities: A
Deniz Oke1, Ozge Gulsum Illeez2, Esra Giray2
1Department of Physical Medicine and Rehabilitation, Health Sciences University Gaziosmanpasa Training and Research Hospital, Istanbul 34255, Turkey.
Abstract:
Background: Large language models (LLMs) such as DeepSeek-V3, Google Gemini 2.0 Pro, and ChatGPT-4o are increasingly used by patients seeking online medical information. However, their accuracy, reliability, and reproducibility in patient-facing content related to chest wall deformities (CWD) remain unclear. This study aimed to compare the performance of three contemporary LLMs in generating information on pectus excavatum, pectus carinatum, and related thoracic deformities. Methods: Eighty patient-facing questions were developed across eight thematic domains and independently submitted to each model using newly created accounts over two consecutive days. Accuracy was assessed using a validated four-point rubric by blinded physiatrists, and reproducibility was evaluated using agreement metrics and weighted Cohen's kappa. Results: ChatGPT-4o achieved the highest overall accuracy (median score: 1.20), the greatest proportion of fully accurate responses, and the lowest hallucination rate (5.0%). Gemini showed intermediate accuracy, while DeepSeek-V3 demonstrated the lowest accuracy and highest hallucination rate (11.25%). Across all models, general-information and quality-of-life domains had the best performance, whereas treatment-related questions showed the most errors. Reproducibility was highest for ChatGPT-4o (weighted κ = almost perfect), followed by Gemini and DeepSeek-V3. Inter-rater reliability was substantial (Fleiss' κ = 0.69). Conclusions: Contemporary LLMs can generate largely accurate and reproducible patient-facing information on CWD, with ChatGPT-4o showing the strongest overall performance. This study provides the first domain-specific comparative evaluation of LLMs in CWD and integrates reproducibility analysis alongside accuracy and reliability assessment. While these tools may support patient education, treatment-related responses require caution, and LLMs should be used as adjuncts rather than substitutes for clinical counseling.
More Related Videos
06:18Pedicle Screw Placement Using an Augmented Reality Head-Mounted Display in a Porcine Model
Published on: May 24, 2024
05:56Implementation of Non-invasive Point of Care Transient Elastography for Evaluation of Liver Disease in Pediatric Populations with Cystic Fibrosis
Published on: August 29, 2025
Related Concept Videos
Pneumothorax-II
Clinical Manifestations:
Radiological Investigation III: Pulmonary Angiogram and PET Scan
Pulmonary Angiogram
A Pulmonary Angiogram is an invasive procedure involving injecting a contrast medium through a catheter threaded into the pulmonary artery or the right side of the heart to visualize the pulmonary vasculature. Computed Tomography (CT) scans have mainly replaced this...
Endoscopic Studies II: Thoracocentesis
Description
Excess pleural fluid or air may accumulate in some respiratory disorders in the thoracic cavity. To treat pleural effusion, a physician conducts thoracentesis by carefully piercing the chest wall and entering...