Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 14, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
06:19

Constructing and Visualizing Models using Mime-based Machine-learning Framework

Published on: July 22, 2025

A framework for efficient scientific diagram captioning using mixture-of-experts and low-rank adaptation.

Deepika Kamboj1,2, Gaurav Harit3

  • 1Computer Science and Engineering, IIT Jodhpur, Jodhpur, Rajasthan, India. kamboj.5@iitj.ac.in.

Scientific Reports
|May 11, 2026
PubMed
Summary

Related Concept Videos

Stereotype Content Model02:16

Stereotype Content Model

The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence categorization, a person will feel...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Representation learning approach for understanding structured documents.

Scientific reports·2025
Same author

An interactive medical image segmentation framework using iterative refinement.

Computers in biology and medicine·2017
See all related articles

This study introduces an efficient framework for generating scientific diagram captions using a Mixture of Experts (MoE) and Low-Rank Adaptation (LoRA). The model achieves high-quality captions with faster inference times, outperforming existing vision-language models.

Area of Science:

  • Scientific visualization
  • Artificial intelligence
  • Multimodal learning

Background:

  • Diagram captioning is complex due to intricate structures and domain-specific content.
  • Existing vision-language models (e.g., BLIP-2, MiniGPT-4, LLaVA) are inefficient for specialized diagram captioning tasks.
  • A need exists for resource-efficient and high-quality diagram captioning solutions.

Purpose of the Study:

  • To develop a novel framework for efficient and accurate scientific diagram caption generation.
  • To improve upon the performance of general-purpose vision-language models in a specialized domain.
  • To introduce a new benchmark dataset and evaluation for diagram captioning.

Main Methods:

  • Proposed a framework integrating Mixture of Experts (MoE) with Low-Rank Adaptation (LoRA).
Keywords:
AI2D-caption datasetDiagram captioningLow-rank adaptation (LoRA)Mixture of experts (MoE)Multimodal learning

Related Experiment Videos

Last Updated: Jun 14, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
06:19

Constructing and Visualizing Models using Mime-based Machine-learning Framework

Published on: July 22, 2025

  • Trained the model on the AI2D-Caption dataset, featuring expert-annotated scientific diagrams.
  • Employed a two-stage approach: pretraining for multimodal alignment and fine-tuning with domain-specific experts and LoRA.
  • Utilized a gating network to dynamically route representations to relevant experts.
  • Main Results:

    • Achieved superior performance compared to zero-shot general vision-language models.
    • Obtained an average BLEU score of 0.42 and an average SPICE score of 0.38.
    • Maintained an efficient average inference time of 3 seconds.
    • Demonstrated superior performance with fewer parameters than competing models.

    Conclusions:

    • The proposed MoE with LoRA framework offers a resource-efficient and effective solution for scientific diagram captioning.
    • The model significantly enhances caption quality and inference speed for specialized visual data.
    • This work contributes a new benchmark and a high-performing model for the diagram captioning task.