Related Experiment Video
Updated: Jun 13, 2026

Electromagnetic Navigation Transthoracic Nodule Localization for Minimally Invasive Thoracic Surgery
Published on: May 4, 2022
Performance of Multimodal Large Language Models in Detection and Position Assessment of Thoracic Devices on Chest
Hamza Eren Güzel1, Cemre Özenbaş2, Babak Saravi3
1Department of Radiology, İzmir City Hospital, University of Health Sciences, İzmir 35540, Türkiye.
Abstract:
Background: Accurate identification and positioning of thoracic devices on chest radiographs is critical for patient safety in intensive care. Multimodal large language models (LLMs) offer potentially generalizable automated evaluation, but their performance in this domain is underexplored. Methods: Three multimodal LLMs (GPT-4o, gpt-4o-2024-08-06; Gemini 3.1 Flash Lite Preview; Claude Sonnet 4.6) were evaluated on 4813 chest radiographs from the RANZCR CLiP dataset for device presence and positioning of ETT, NGT, CVC, and Swan-Ganz catheters. Performance was quantified with 95% Wilson confidence intervals, balanced accuracy, MCC, Cochran's Q, Bonferroni-corrected McNemar, and Cohen's/Fleiss' kappa. Six additional analyses were performed: a blinded paired reader study (n = 377; two board-certified radiologists, blinded to ground truth and to all LLM outputs), external validation on PadChest (n = 200, device-presence detection only-PadChest lacks granular position labels), three-variant prompt-sensitivity analysis (n = 103), repeat-inference stability across three runs (n = 50), systematic error taxonomy, and a failure-case analysis. Results: Device-presence performance varied widely across models; abnormal-position sensitivity was uniformly poor (MCC ≤ 0.028; balanced accuracy 0.41-0.53). Inter-model agreement was poor to slight (Fleiss' κ: 0.005-0.383 for presence; -0.280 to -0.025 for classification). Radiologists numerically outperformed all three LLMs in 42/42 paired comparisons; the superiority was statistically significant after Bonferroni correction in 33/42 (32/42 at p < 0.001). PadChest replicated the negative finding for device-presence detection (malposition not externally validated). Prompts and inference stochasticity introduced 2-3× sensitivity swings and run-to-run κ from 0.20 to 0.85. Case failures concentrated systematically in multi-device cases (p < 0.0001) but not in abnormal-position cases (p = 0.14). Conclusions: Current general-purpose multimodal LLMs are not yet reliable for autonomous thoracic-device assessment; their failure patterns are structurally characterizable across models, prompts, and case types and support, at most a circumscribed role, as adjunct device-presence screening tools. The findings do not generalize to purpose-built, regulator-approved clinical AI systems.
Related Concept Videos
Assessment of Airway, Skin Color, and Use of Accessory Muscles
Introduction
The initial evaluation of a patient's respiratory system...
Radiological Investigation III: Pulmonary Angiogram and PET Scan
Pulmonary Angiogram
A Pulmonary Angiogram is an invasive procedure involving injecting a contrast medium through a catheter threaded into the pulmonary artery or the right side of the heart to visualize the pulmonary vasculature. Computed Tomography (CT) scans have mainly replaced this...
