Related Experiment Video
Updated: Jun 7, 2026

A Multimodal Imaging Framework to Advance Phenotyping of Living Label-free Breast Cancer Cells
Published on: August 22, 2025
Application of multimodal large language models for HER2/CEN17 FISH signal detection
Kamil Stanisław Fryszkowski1, Tomasz Les1
1Warsaw University of Technology, Plac Politechniki 1, Warsaw, 00-661, Poland.
Background And Objective:
Fluorescence in situ hybridisation (FISH) is a reference technique for HER2 gene amplification assessment, yet manual signal counting is labour-intensive and subject to inter-observer variability. This study evaluates the zero and few-shot capability of multimodal large language models for automatic counting of HER2 and CEN17 signals in single-nucleus images.
Methods:
A data set of 240 nuclei, categorised by difficulty (simple, complex, and cluster), was extracted from clinical slides acquired at 100× magnification. Three models-Gemini 2.5 Pro, GPT-5, and Claude Sonnet 4.5 were tested using natural language prompts. In the main evaluation setting on original colour images, Gemini 2.5 Pro outperformed the others, achieving a mean absolute error (MAE) of 0.47 signals for simple nuclei and 0.92 for complex cases.
Results:
The model demonstrated high consistency across repeated runs (median absolute deviation of 0.22) and relied heavily on colour information, as the aggregated MAE worsened to 1.60 on greyscale inputs. Bland-Altman analysis revealed a systematic overcounting bias of 0.39 signals (p<0.001), driven primarily by HER2 (bias 0.75, p<0.001), whereas CEN17 counts showed no significant bias (bias 0.09, p=0.213).
Conclusions:
These results indicate that the model is more likely to count background features as additional signals than to miss true signals. Although current error rates suggest they are not yet sufficiently reliable for autonomous clinical decision-making, the results demonstrate that general-purpose multimodal models can achieve sub-signal accuracy without specific training, indicating that this direction is worth pursuing. This work is intended as an exploratory feasibility study examining the reasoning capabilities of MLLMs as zero-shot counting engines.
More Related Videos
09:44Recording and Analyzing Multimodal Large-Scale Neuronal Ensemble Dynamics on CMOS-Integrated High-Density Microelectrode Array
Published on: March 8, 2024
06:51Dual-modality Molecular Cartography: Integrating Multiplex mRNA Detection with Protein Imaging Mass Cytometry
Published on: November 14, 2025