Related Experiment Video
Updated: Feb 13, 2026

Author Spotlight: Advancing 3D Modeling for Enhanced Diagnosis and Treatment of Pulmonary Nodules in Early-Stage Lung Cancer
Published on: October 13, 2023
A comparative accuracy study of multimodal LLMs, VLM and agent-based framework for pulmonary nodule detection on
Daria Khovanova1, Yuriy Vasilev1,2, Anton Vladzymyrskyy1,3
1Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department, Moscow, Russia.
Background:
Artificial intelligence technologies are being actively introduced in clinical practice. The most promising solutions are AI-assistants based on large language models (LLMs). Determining the feasibility of integrating such applications in clinical practice requires independent performance assessments. This study assessed accuracy of several multimodal LLMs in detecting pulmonary nodules on chest radiographs (CXR).
Methods:
This study included 9 models: Llama 3.2 Vision 90B, Claude 3.5 Sonnet, Claude 3.7 Sonnet, Gemini 2.0 Pro Experimental, Perplexity, CXR-LLaVA, XrayGPT, BiomedCLIP, MedRAX. Each model determined presence or absence of pulmonary nodules in dataset containing 100 CXR, 50 of which contained pulmonary nodules. ROC curves were constructed, diagnostic accuracy metrics were calculated. McNemar's test was used for pairwise accuracy comparisons.
Results:
Best results were achieved by MedRAX framework and BiomedCLIP vision-language model, with accuracy of 0.711 (95% CI 0.613-0.808). Among proprietary single-model LLMs, Claude 3.7 Sonnet demonstrated the best performance: accuracy 0.651 (0.548-0.753). Llama 3.2 Vision 90B, Claude 3.5 Sonnet, Gemini 2.0 Pro Experimental demonstrated matching accuracy values: 0.602 (0.497-0.708).
Conclusion:
MedRAX framework and BiomedCLIP vision-language model showed the highest accuracy values. No statistically significant difference was observed between proprietary and open-source models, which may indicate potential for improving accuracy through refinement of open-source LLM-based models. Overall, accuracy values of evaluated models were insufficient for current clinical practice implementation. These results should be seen as exploratory given the small dataset size, single-centre design, different prompting strategies for foundation and domain-adapted models and use of PNG images instead of DICOM.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Uncertainty in Measurement: Accuracy and Precision
Flail Chest-I
Flail chest is a severe and potentially life-threatening condition characterized by the fracture of three or more adjacent ribs in multiple places. It is most commonly caused by direct impacts and trauma, such as motor vehicle accidents or injuries from a steering wheel impact. It can also occur due to falls in elderly individuals with osteoporosis, or assaults involving sharp objects.
Pathophysiology
The pathophysiology of flail chest is complex, involving fractures of...
Flail Chest-II
Assessment:
1. Clinical Evaluation:
History:
Chest Physiotherapy
Purpose
CPT is primarily used for patients with excessive bronchial secretions who have difficulty clearing...

