Vision language models for scientific image analysis: an evaluation highlighting opportunities and challenges

Prateek Verma1, Minh-Hao Van1, Xintao Wu1

  • 1Department of Electrical Engineering and Computer Science, University of Arkansas, Fayetteville, AR USA.

Npj Computational Materials
|June 25, 2026
PubMed
Summary

Vision language models (VLMs) show promise for analyzing scientific microscopy images in tasks like classification and segmentation. While not yet expert-level, models like ChatGPT and Gemini demonstrate improved comprehension and segmentation capabilities.