Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative Accuracy, Stability, and Correctability of Large Language Models in Otolaryngology and Pharmacovigilance
Filippo Bruno1, Lise Sogalow1, Bertrand Blankert2
1Department of Surgery, Research Institute for Language Science and Technology, University of Mons, Mons, Belgium.
Objective:
To compare the clinical and pharmacovigilance performance, stability, and correctability of 3 large language models (LLMs) in otolaryngology outpatient care.
Study Design:
Prospective case series.
Setting:
Multicenter University Hospitals.
Methods:
Consecutive adults (August-October 2024) with established primary diagnoses were entered into ChatGPT-4o, Gemini-1.5-Pro, and Claude-3.5-Sonnet using only history and physical examination findings (no complementary tests) via standardized prompts. Two blinded otolaryngologists rated clinical accuracy with the Artificial Intelligence Performance Instrument (AIPI); 2 blinded pharmacists rated pharmacological information on a 5-point Likert scale. Errors were fed back to models and all cases were re-queried one month later. Interrater reliability used ICC; stability used Cronbach's α. Group differences used Kruskal-Wallis.
Results:
Fifty-one patients with 60 diagnoses across otolaryngology subspecialties were consecutively recruited (38 females (74.5%); mean age of 42.4 ± 17.4 years). All LLMs recommended significantly more additional examinations than practitioners (P = .001), with a significant increase of the number of recommended additional examinations after regenerated inputs for ChatGPT-4o and Claude-3.5-Sonnet, respectively. Claude-3.5-Sonnet and ChatGPT-4o outperformed Gemini-1.5-Pro for AIPI-clinical management (P = .001) and pharmacovigilance findings (P = .001). The physicians (ICC = 0.853) and the pharmacists (ICC = 0.991) demonstrated an almost perfect interrater reliability. All LLMs demonstrated an almost perfect clinical stability (α = 0.831-0.856), though human feedback did not significantly reduce misdiagnosis rates in subsequent interactions.
Conclusion:
In outpatient ENT cases using clinical features alone, ChatGPT-4o and Claude-3.5-Sonnet deliver higher clinical and pharmacovigilance performance than Gemini-1.5-Pro, with almost perfect interrater reliability and stable outputs. Re-querying after feedback did not improve accuracy, questioning short-term correctability.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
04:04Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Analysis of Population Pharmacokinetic Data
Pharmacokinetic Models: Overview
There are three primary types of models: empirical, compartment, and physiological. Empirical models, with minimal...