Related Experiment Video
Updated: Jul 18, 2026

04:48
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
ViT-MultiRAGNet: A scalable and reliable retrieval-augmented Vision Transformer framework for memory-guided feature
N Thirupathi Rao1, Ch V V Ramana2, Faten Khalid Karim3
1Department of Computer Science and Engineering, Vignan's Institute of Information Technology (A), Visakhapatnam, Andhra Pradesh, India.
Plos One
|July 15, 2026
Summary
This study introduces ViT-MultiRAGNet, a novel deep learning framework that enhances breast cancer diagnosis by integrating mammographic features with historical case data. The system significantly improves diagnostic accuracy and interpretability for clinical decision support.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Deep Learning for Breast Cancer Detection
- Computer-Aided Diagnosis Systems
Background:
- Mammography diagnosis is challenged by image variability and limited use of historical patient data.
- Current deep learning models often neglect historical case context, impacting diagnostic accuracy.
- ViT-MultiRAGNet integrates multi-view mammographic features with historical case evidence for improved inference.
Purpose of the Study:
- To develop a retrieval-augmented Vision Transformer framework (ViT-MultiRAGNet) for enhanced breast cancer diagnosis.
- To leverage multi-view 2D mammographic features and historical case data for improved diagnostic accuracy.
- To create an evidence-based decision support system for breast cancer screening and diagnosis.
Main Methods:
- Utilized a Vision Transformer encoder for multi-view mammogram feature extraction.
- Employed a retrieval-augmented memory bank to identify and incorporate evidence from similar historical cases.
- Integrated global mammographic features, lesion representations, and retrieved historical evidence via a Retrieval-Augmented Generation (RAG) module and multi-head cross-attention.
Main Results:
- ViT-MultiRAGNet achieved high accuracy (0.978±0.009 on RTM, 0.961±0.011 on CBIS-DDSM) and AUC-ROC (0.998±0.009 on RTM, 0.989±0.010 on CBIS-DDSM).
- Segmentation evaluation showed improved edge alignment with Dice scores of 0.882 (RTM) and 0.795 (CBIS-DDSM).
- The RAG mechanism enhanced retrieval-guided fusion, maintaining efficient inference at 0.31 seconds per image.
Conclusions:
- The ViT-MultiRAGNet framework significantly improves diagnostic performance, robustness, and clinical interpretability in breast cancer detection.
- The integration of Vision Transformer-based features, retrieval-augmented fusion, and memory-guided evidence enhances diagnostic capabilities.
- This approach offers a transparent, evidence-based decision support system for clinicians in breast cancer screening and diagnosis.