Related Experiment Video
Updated: Mar 21, 2026

11:19
Multimodal Hierarchical Imaging of Serial Sections for Finding Specific Cellular Targets within Large Volumes
Published on: March 20, 2018
11.0K
Visual recognition limitations in multimodal large language models: A comparative analysis of histological image
Volodymyr Mavrych1, Einas M Yousef1, Ahmed Yaqinuddin1
1College of Medicine, Alfaisal University, Riyadh, Kingdom of Saudi Arabia.
PLOS Digital Health
|March 19, 2026
Summary
Gemini 2.5 Flash demonstrated superior histological image analysis compared to GPT-4o, Claude Sonnet 4, and Copilot. This study highlights current limitations in multimodal large language models' visual recognition for medical imaging.
Area of Science:
- Artificial Intelligence in Medicine
- Digital Pathology
- Histology Image Analysis
Background:
- Multimodal large language models (LLMs) show promise for medical image analysis.
- Their performance in specialized fields like histology is largely unknown.
- This study evaluates leading LLMs for histological image interpretation.
Purpose of the Study:
- To systematically assess the performance of four leading multimodal LLMs in histological image interpretation.
- To evaluate their visual recognition capabilities in a specialized medical domain.
- To establish benchmarks for multimodal LLM development in medical imaging.
Main Methods:
- Four multimodal LLMs (Gemini 2.5 Flash, GPT-4o, Copilot, Claude Sonnet 4) were tested.
- 144 histological images across four tissue types and three magnifications were used.
- Expert faculty graded LLM responses on tissue identification, morphology, and function using a 4-point scale.
Main Results:
- Gemini 2.5 Flash achieved the highest performance (3.35/4.00), significantly outperforming others.
- Copilot and GPT-4o tied for second (2.76/4.00), with Claude Sonnet 4 lowest (2.55/4.00).
- Performance varied by tissue type, with epithelial tissue showing the most inter-model variation.
Conclusions:
- Gemini 2.5 Flash excels in histological image analysis among current multimodal LLMs.
- Significant gaps exist between text and visual processing in LLMs, indicating architectural constraints.
- Further innovation in specialized visual processing is needed for advanced medical imaging applications.

