Related Experiment Video
Updated: Aug 6, 2026

Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
Comparative Evaluation of ChatGPT and Gemini in Approximating Clinically Confirmed Diagnoses From Structured
Sedra Arnous1, Rama Mdaghmesh1, Alaa Rajab1
1Faculty of Medicine, Department of Medicine and Surgery, Syrian Private University, Damascus, SYR.
Introduction:
The utilization of large language models (LLMs) in healthcare shows promising potential for addressing the global shortage of audiologists. However, their effectiveness in understanding and interpreting complex, live, real-world structured audiological data requires further investigation, as previous studies have primarily focused on theoretical situations and scenarios. This study aimed to assess and compare the diagnostic inference abilities of ChatGPT (OpenAI, San Francisco, USA) and Gemini (Google, Mountain View, USA) in differentiating clinically confirmed diagnoses based on numerical pure-tone audiogram data.
Methods:
This prospective comparative study was conducted at Damascus Hospital (Al-Mujtahid Hospital), Damascus, Syria. Pure-tone audiograms were collected from 309 patients aged between six and 90 years in the audiology department. Audiograms were converted into a numerical format and independently analyzed by ChatGPT-4o and Gemini 1.5 Pro. The outputs were then compared with the reference standard, namely, expert clinical diagnosis, to evaluate accuracy, specificity, sensitivity, and Cohen's kappa for both models.
Results:
Hearing loss was identified in 274 of 309 patients (88.7%). ChatGPT demonstrated 60.8% accuracy, with a weighted sensitivity of 68.9%, a weighted specificity of 94.5%, and a Cohen's kappa of 0.578. Gemini achieved a significantly higher accuracy of 86.1%, with a weighted sensitivity of 87.7%, a weighted specificity of 97.5%, and a Cohen's kappa of 0.848. Both models demonstrated high sensitivity and specificity for certain conditions, such as otosclerosis and presbycusis. The most common diagnosis was otosclerosis (27%), followed by presbycusis (16%).
Conclusion:
The findings suggest that Gemini 1.5 Pro demonstrated stronger performance, achieving an accuracy of 86.1%, whereas ChatGPT-4o achieved 60.8% in interpreting real-world pure-tone audiograms. Both models showed high agreement in common audiometric patterns. These models should not be relied upon independently but may serve as supportive tools under clinician supervision to reduce diagnostic errors.