Related Experiment Video
Updated: Apr 18, 2026

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
Published on: October 27, 2023
MILU: a consensus ensemble benchmark for multimodal medical imaging lecture understanding.
Md Motaleb Hossen Manik1, Md Zabirul Islam1, Ge Wang2
1Rensselaer Polytechnic Institute, Department of Computer Science, Troy, New York, United States.
Vision-language models (VLMs) show high formatting reliability but low semantic consistency on scientific lecture slides. The Medical Imaging Lecture Understanding (MILU) benchmark reveals significant variability in their structured understanding.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Medical Education
Background:
- Vision-language models (VLMs) are increasingly utilized for interpreting multimodal educational content.
- The reliability of VLMs on complex scientific lecture slides with diagrams, equations, and dense text is not well-established.
Purpose of the Study:
- Introduce the Medical Imaging Lecture Understanding (MILU) benchmark, a large-scale dataset for evaluating VLM performance on medical imaging lectures.
- Characterize cross-model variability in the structured understanding of scientific lecture slides by VLMs.
Main Methods:
- Evaluated four prominent VLMs (LLaVA-OneVision, InternVL3-14B, Qwen2-VL-7B, Qwen3-VL-4B) using unified prompts to generate structured JSON outputs.
- Assessed parsing coverage, pairwise semantic agreement, lecture-level patterns, and alignment with a consensus ensemble.
Main Results:
- All evaluated VLMs demonstrated high JSON formatting coverage (92%-99%) but exhibited extremely low semantic agreement (pairwise Jaccard indices 0.03-0.09).
- Lecture-level stability varied, with mathematically structured content showing higher consistency than diagram-heavy slides.
- A consensus ensemble showed modest alignment with individual models, highlighting both areas of convergence and systematic disagreement.
Conclusions:
- MILU serves as the inaugural benchmark for assessing structured understanding of scientific lecture slides.
- Current VLMs excel in formatting but lack semantic consistency, indicating a need for improved scientific lecture interpretation methods.
- This benchmark lays the groundwork for future expert-annotated datasets, diagram/math-aware VLM development, and enhanced scientific content analysis.
Related Concept Videos
Imaging Studies I: CT and MRI
Description of the Procedures
Computed Tomography (CT) scan:
Computed Tomography (CT) scans use X-ray technology to generate detailed images of bones, organs, and tissues. During the scan, the patient lies on a moving table...
Imaging Studies III: Computed Tomography
Imaging Studies VII: Vascular Imaging
Imaging Studies for Cardiovascular System V: CT
Imaging Studies for Cardiovascular System IV: CMRI
Imaging Studies IV: Magnetic Resonance Imaging