Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 26, 2026

Functional Magnetic Resonance Imaging (fMRI) of the Visual Cortex with Wide-View Retinotopic Stimulation
07:11

Functional Magnetic Resonance Imaging (fMRI) of the Visual Cortex with Wide-View Retinotopic Stimulation

Published on: December 8, 2023

Vision language models for scientific image analysis: an evaluation highlighting opportunities and challenges.

Prateek Verma1, Minh-Hao Van1, Xintao Wu1

  • 1Department of Electrical Engineering and Computer Science, University of Arkansas, Fayetteville, AR USA.

Npj Computational Materials
|June 25, 2026
PubMed
Summary

Vision language models (VLMs) show promise for analyzing scientific microscopy images in tasks like classification and segmentation. While not yet expert-level, models like ChatGPT and Gemini demonstrate improved comprehension and segmentation capabilities.

Related Concept Videos

Vision01:24

Vision

Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

High-temperature electrical performance and interface adhesion performance of SiC-based Pt thin-film RTDs.

Microsystems & nanoengineering·2026
Same author

Fairness across domains: a unified fairness-aware framework for domain generalization and unsupervised adaptation.

Frontiers in big data·2026
Same author

Multi-Modal Foundation Models for Computational Pathology: A Survey.

Transactions on machine learning research·2026
Same author

Histopathology-centered Computational Evolution of Spatial Omics: Integration, Mapping, and Foundation Models.

ArXiv·2026
Same author

Lithiated PAA-Coated SiO<sub>x</sub> Anode for Stable and High-Capacity Lithium-Ion Batteries: Interfacial Regulation and Volume Expansion Suppression.

Small (Weinheim an der Bergstrasse, Germany)·2026
Same author

Coexistence rules for small, antagonistically interacting microbial communities.

PLoS computational biology·2025

Area of Science:

  • Scientific image analysis
  • Microscopy
  • Artificial intelligence in science

Background:

  • Vision language models (VLMs) like ChatGPT, Gemini, Llama, and LLaVA excel at processing visual and textual data.
  • The Segment Anything Model (SAM) demonstrates advanced image segmentation.
  • Microscopy images are crucial in biology, medicine, and materials science.

Purpose of the Study:

  • To evaluate the performance of advanced VLMs (ChatGPT-5, Gemini-2.5Pro, Llama-3.2V, LLaVA-1.5) and SAM (SAM-2) on microscopy image analysis tasks.
  • Assess capabilities in classification, segmentation, counting, and visual question answering (VQA).

Main Methods:

  • Utilized microscopy images for evaluation.
  • Tested models including ChatGPT-5, Gemini-2.5Pro, Llama-3.2V, LLaVA-1.5, and SAM-2.
Keywords:
Electrical and electronic engineeringMaterials science

Related Experiment Videos

Last Updated: Jun 26, 2026

Functional Magnetic Resonance Imaging (fMRI) of the Visual Cortex with Wide-View Retinotopic Stimulation
07:11

Functional Magnetic Resonance Imaging (fMRI) of the Visual Cortex with Wide-View Retinotopic Stimulation

Published on: December 8, 2023

  • Focused on tasks: classification, segmentation, counting, and VQA.
  • Main Results:

    • ChatGPT and Gemini showed strong performance in image comprehension.
    • SAM demonstrated excellent object isolation capabilities.
    • All models showed improvement over previous versions but did not reach domain expert accuracy, especially with complex image artifacts.

    Conclusions:

    • VLMs show significant potential for advancing scientific image analysis.
    • Further development is needed to achieve expert-level performance on complex microscopy data.
    • These models offer a promising foundation for future scientific discovery tools.