Related Experiment Video
Updated: Sep 8, 2025

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
565
Performance of generative AI across ENT tasks: A systematic review and meta-analysis
Sholem Hack1, Rebecca Attal1, Armin Farzad1
1City St. Georges University London School of Medicine, Program Delivered by University of Nicosia at the Chaim Sheba Medical Center, Ramat Gan, Israel.
Auris, Nasus, Larynx
|September 5, 2025
Summary
Generative AI shows promise in otolaryngology education and communication tasks. However, its inconsistent performance in clinical reasoning necessitates further research for safe integration into practice.
Area of Science:
- Otolaryngology
- Artificial Intelligence
- Medical Informatics
Background:
- Generative AI, particularly Large Language Models (LLMs), is increasingly explored for medical applications.
- Evaluating the efficacy of LLMs in specialized fields like otolaryngology is crucial for understanding their potential and limitations.
Purpose of the Study:
- To systematically assess the diagnostic accuracy, educational utility, and communication capabilities of generative AI in otolaryngology.
- To analyze the performance variations of different LLM versions (e.g., GPT-4 vs. GPT-3.5) in otolaryngology-related tasks.
Main Methods:
- A comprehensive literature search was conducted across major scientific databases (PubMed, Embase, Scopus, Web of Science, IEEE Xplore) for studies published between January 2022 and March 2025.
- Eligible studies evaluating text-based generative AI in otolaryngology were screened and assessed using JBI and QUADAS-2 tools.
- A random-effects meta-analysis was performed on diagnostic accuracy data, with subgroup analyses based on task type and AI model version.
Main Results:
- Ninety-one studies were included, with 43 providing quantitative diagnostic accuracy data across 59 model-task pairs.
- Pooled diagnostic accuracy for generative AI in otolaryngology was 72.7%, with higher accuracy in educational (83.0%) and diagnostic imaging (84.9%) tasks compared to clinical decision support (67.1%).
- GPT-4 demonstrated superior performance over GPT-3.5, though challenges like hallucinations and performance variability were observed in complex clinical reasoning.
Conclusions:
- Generative AI demonstrates strong performance in structured otolaryngology tasks, particularly in education and communication.
- The current inconsistent performance in clinical reasoning limits the standalone application of generative AI in patient care.
- Future research should prioritize mitigating AI hallucinations, establishing standardized evaluation metrics, and conducting prospective validation studies for safe clinical integration.
