Related Experiment Video
Updated: Jul 9, 2026

Rapid Identification of Gram Negative Bacteria from Blood Culture Broth Using MALDI-TOF Mass Spectrometry
Published on: May 28, 2014
From Gram Stain to Decision Support: Performance of Multimodal Large Language Models in Blood Culture Microscopy
Canset Nur Aydogan1, Ilgaz Kazaz2, Tuğrul Hoşbul3
1Department of Medical Microbiology, Ordu State Hospital, 52200, Ordu, Turkey. aydogancanset@gmail.com.
Multimodal large language models (LLMs) show promise for Gram stain interpretation in clinical microbiology, achieving high accuracy for Gram type but lower accuracy for morphology. LLMs are best used as expert-supervised decision-support tools, not standalone systems.
Area of Science:
- Medical Microbiology
- Artificial Intelligence in Diagnostics
- Computational Pathology
Background:
- Multimodal large language models (LLMs) present potential for image-based diagnostics.
- Reliability of LLMs in routine clinical microbiology, specifically Gram stain interpretation, is not well-established.
- Gram staining remains a cornerstone of bacterial identification in clinical laboratories.
Purpose of the Study:
- To evaluate the diagnostic accuracy of three leading multimodal LLMs (ChatGPT-4o, Gemini 2.5 Flash, Claude Opus) in interpreting Gram-stained blood culture smears.
- To benchmark LLM performance against expert medical microbiologists and automated identification systems.
- To assess LLM consistency and identify areas of strength and limitation in Gram stain interpretation.
Main Methods:
- Prospective diagnostic accuracy study involving 100 Gram-stained blood culture smear images.
- Evaluation of LLMs using standardized zero-shot prompts across two independent runs for intramodel consistency.
- Benchmarking against three expert medical microbiologists and automated organism identification as the reference standard.
Main Results:
- Expert microbiologists achieved 100% concordance with the reference standard for Gram type and major morphology.
- LLMs demonstrated high accuracy for Gram-type classification (95-98%) but lower accuracy for cellular morphology (84-85%), resulting in combined accuracies of 82-84%.
- Performance varied by organism class, with lower accuracy for Gram-positive bacilli and yeast; ChatGPT and Gemini outperformed Claude on fine-grained morphology.
Conclusions:
- Multimodal LLMs exhibit promising baseline performance for Gram-stained blood culture interpretation, particularly for Gram type and major morphology.
- Current limitations include reduced accuracy in combined interpretation, fine-grained morphology, and less common organism classes, indicating out-of-the-box models are insufficient for standalone use.
- Findings support the use of LLMs as expert-supervised decision-support tools, emphasizing the need for task-specific optimization and multicenter validation prior to clinical implementation.
More Related Videos
09:07Direct Microbial Identification using An Automated Microbial Identification System to Facilitate the EUCAST RAST Method Without Mass Spectrometry
Published on: May 24, 2024
11:25Preparation of a Blood Culture Pellet for Rapid Bacterial Identification and Antibiotic Susceptibility Testing
Published on: October 15, 2014
Related Concept Videos
Sputum Studies I: Gram Stain, cytology, and Acid-fast smear and culture
Gram Stain
The Gram Stain is an integral part of sputum studies. It involves the staining of sputum, which permits...
Special Staining Techniques
Differential Staining Technique
Fixation and Sectioning
The simplest type of preparation is the wet mount, in which the specimen is placed in a drop of liquid on the slide. A liquid specimen can be directly deposited on the slide using a dropper. Solid specimens, such as skin scraping, can be placed on the slide before adding a drop of liquid to prepare the wet mount. Sometimes the liquid is simply water, but stains are often added...
Methods to Assess Microbial Populations
Automated Microbial Diagnostics