Related Experiment Video
Updated: Mar 16, 2026

A Rapid Method for Multispectral Fluorescence Imaging of Frozen Tissue Sections
Published on: March 30, 2020
Exploring clinical feasibility of zero-shot learning in a large language model for immunofixation electrophoresis
1Department of Biochemistry, School of Medicine, Kocaeli University, Kocaeli, Turkey.
Insights
Large language models (LLMs) show promise for interpreting immunofixation electrophoresis (IFE) images. Gemini demonstrated superior performance over ChatGPT in automated IFE analysis, though further validation is needed.
Area of Science:
- Artificial Intelligence
- Medical Diagnostics
- Laboratory Medicine
Background:
- Immunofixation electrophoresis (IFE) is the standard for detecting monoclonal immunoglobulins but is subjective.
- Multimodal large language models (LLMs) offer zero-shot visual reasoning capabilities.
- Automating IFE interpretation could reduce variability and improve efficiency.
Purpose of the Study:
- To evaluate the feasibility and diagnostic performance of zero-shot multimodal LLMs for IFE image interpretation.
- To compare the performance of different LLMs in classifying IFE patterns.
- To assess the impact of prompt design on LLM performance.
Main Methods:
- A dataset of 487 IFE images across ten classes was used.
- Two multimodal LLMs (ChatGPT-5.2, Gemini3 Pro) were evaluated using a zero-shot visual question answering (ZS-VQA) framework.
- Performance was assessed using precision, recall, F1-score, and confusion matrix analysis with simple and detailed prompts.
Main Results:
- Gemini outperformed ChatGPT, achieving high F1-scores for clinically relevant monoclonal categories (e.g., MK, AK, GK).
- Detailed prompts significantly improved recall and F1-scores for both models.
- Both models struggled with visually subtle patterns, indicating areas for improvement.
Conclusions:
- Zero-shot multimodal LLMs, especially Gemini, show potential for automated IFE interpretation.
- Further optimization and validation are necessary for clinical implementation due to performance variability.
- Prompt engineering is crucial for enhancing LLM diagnostic accuracy in IFE.
Introduction:
Immunofixation electrophoresis (IFE) is the gold standard method for detecting and typing monoclonal immunoglobulins, but its interpretation requires expert evaluation and remains susceptible to subjectivity and variability. Recent advances in multimodal large language models (LLMs) have enabled zero-shot visual reasoning without task-specific training. This study aimed to evaluate the feasibility and diagnostic performance of zero-shot multimodal LLMs for automated interpretation of IFE images.
Materials And Methods:
A dataset of 487 immunofixation electrophoresis images representing ten semantic classes (AK, AL, GK, GL, K, L, LGL, MK, ML, and NONE) was retrospectively collected and annotated by expert laboratory specialists. Two multimodal LLMs, ChatGPT-5.2 and Gemini3 Pro, were evaluated using a zero-shot visual question answering (ZS-VQA) framework without additional training or fine-tuning. Each image was analyzed using two prompt configurations: a simple prompt and a detailed prompt with structured diagnostic guidance. Model performance was assessed using precision, recall, F1-score, and confusion matrix analysis.
Results:
Gemini outperformed ChatGPT across most classes, particularly with detailed prompts, achieving high F1-scores in clinically relevant monoclonal categories such as MK (89.80), AK (84.75), and GK (71.54). Detailed prompting consistently improved recall and F1-scores, underscoring the importance of prompt design. In contrast, ChatGPT showed lower recall, F1-scores, and classification consistency. Both models performed less effectively in visually subtle patterns.
Conclusions:
Zero-shot multimodal LLMs, particularly Gemini, show promising potential for interpreting IFE images without task-specific training. However, performance variability and limitations in certain classes indicate that further optimization and validation are required before clinical implementation.

