Related Experiment Video
Updated: May 22, 2026

06:13
Endoscopic Septoplasty with Limited Two-line Resection: Minimally Invasive Surgery for Septal Deviation
Published on: June 20, 2018
Reliability of Multimodal LLMs for Sinusitis with Polyps vs. Sinusitis Without Polyps Classification from Paranasal
Buket Yağcı1, Sergen Palaz2, Rezarta Taga Senirli3
1Department of Radiology, Antalya Training and Research Hospital, 07100, Antalya, Turkey. buketyagci@hotmail.com.
Journal of Imaging Informatics in Medicine
|May 20, 2026
Summary
ChatGPT-4o showed the highest accuracy in distinguishing sinusitis with nasal polyps from CT scans. This study highlights the varying diagnostic reliability of multimodal large language models (LLMs) for medical imaging analysis.
Area of Science:
- Medical Imaging AI
- Artificial Intelligence in Radiology
- Computational Pathology
Background:
- Distinguishing sinusitis with nasal polyps from sinusitis without polyps is crucial for accurate diagnosis and treatment.
- Computed tomography (CT) is a primary imaging modality for evaluating paranasal sinus disease.
- Large Language Models (LLMs) are emerging as tools for medical image analysis, but their diagnostic reliability requires thorough evaluation.
Purpose of the Study:
- To compare the diagnostic reliability of commercially available multimodal large language models (LLMs) in differentiating sinusitis with nasal polyps from sinusitis without polyps using paranasal sinus CT images.
- To assess the performance of different LLM versions and vendors in a specific clinical diagnostic task.
Main Methods:
- A retrospective study analyzed 175 adult paranasal sinus CT scans (80 with polyps, 95 without).
- Four representative anonymized slices per case were presented to three multimodal LLMs (ChatGPT-4o, ChatGPT-5, Gemini 2.5 Pro) using a zero-shot approach.
- Model outputs were evaluated by expert radiologists, and performance metrics (accuracy, sensitivity, specificity, kappa) were calculated.
Main Results:
- ChatGPT-4o achieved the highest diagnostic accuracy (0.89), sensitivity (0.88), specificity (0.90), and kappa (0.77), significantly outperforming other models.
- ChatGPT-5 demonstrated moderate agreement (accuracy 0.67, kappa 0.33), while Gemini 2.5 Pro showed the lowest performance (accuracy 0.56, kappa 0.11).
- Significant performance differences were observed among multimodal LLMs for CT-based polyp detection.
Conclusions:
- Multimodal LLMs exhibit varying diagnostic reliability for paranasal sinus CT interpretation, with significant differences based on model version and vendor.
- Task-specific validation and continuous monitoring are essential before integrating LLMs into clinical radiology workflows.
- Further research is needed to optimize LLM performance for specific diagnostic tasks in medical imaging.
