Related Experiment Video
Updated: Jul 8, 2026

13:44
Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
43.5K
A Benchmark for Breast Cancer Screening and Diagnosis in Mammogram Visual Question Answering
Jiayi Zhu1, Fuxiang Huang2, Qiong Luo3,4
1Data Science and Analytics Thrust, The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, Guangdong, China.
Nature Communications
|November 27, 2025
Summary
Large vision-language models struggle with mammogram interpretation, performing no better than random guessing. A new dataset and domain-optimized model, LLaVA-Mammo, show significant improvements in mammogram analysis accuracy.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Computer Vision
Background:
- Breast cancer is a leading global malignancy in women.
- Early detection via mammography significantly improves survival rates.
- Large vision-language models (LVLMs) show promise for mammogram analysis, but lack standardized benchmarks for performance comparison.
Purpose of the Study:
- To address the lack of standardized evaluation benchmarks for LVLMs in mammogram interpretation.
- To introduce MammoVQA, a comprehensive dataset for mammogram visual question answering.
- To evaluate the performance of current LVLMs and develop a domain-optimized model for improved mammogram analysis.
Main Methods:
- Creation of MammoVQA dataset by unifying 15 public datasets, including image-level and exam-level data with question-answering pairs.
- Systematic evaluation of 12 high-performance LVLMs (6 general, 6 medical) on the MammoVQA dataset.
- Development and validation of a domain-optimized LLaVA-Mammo model.
Main Results:
- Current high-performance LVLMs demonstrated diagnostic performance equivalent to random guessing on mammogram interpretation tasks.
- The proposed LLaVA-Mammo model achieved significant accuracy improvements: +19.66% in internal validation and +21.21% in external validation compared to the best existing models.
- The MammoVQA dataset provides a standardized benchmark for evaluating LVLMs in mammography.
Conclusions:
- Existing LVLMs are currently unreliable for clinical mammogram interpretation.
- The MammoVQA dataset and the domain-optimized LLaVA-Mammo model represent significant advancements in AI-assisted mammography.
- Further research and development are needed to enhance the reliability and clinical utility of AI models in breast cancer screening.

