Related Experiment Video
Updated: Aug 24, 2026

Assessing Early Stage Open-Angle Glaucoma in Patients by Isolated-Check Visual Evoked Potential
Published on: May 25, 2020
Medrecord-CLIP: enhancing fundus disease diagnosis via EHR-guided vision-language pre-training
Lei Shi1, Wenbin Zhai1, Lei Yu2
1School of Computer and Information, Hefei University of Technology, Hefei, 230009 Anhui China.
None:
Automated analysis of Color Fundus Photography (CFP) is essential for large-scale retinal disease screening. However, conventional vision-only models often rely on rigid categorical labels, neglecting the rich clinical nuances found in medical narratives. While existing Vision-Language Pre-training (VLP) frameworks have explored text-based supervision, they frequently overlook individualized patient contexts within Electronic Health Records (EHRs). To bridge this gap, we construct MedRecordFundus, a large-scale multimodal dataset pairing 21,290 CFP images with expert-curated EHR narratives across four clinical dimensions. Leveraging this resource, we propose MedRecord-CLIP, a knowledge-enhanced foundation model featuring a diagnosis-guided cross-attention mechanism to adaptively extract and fuse salient patient history with diagnostic representations. To improve representation robustness against highly similar clinical descriptions, we introduce a symmetric Multi-Positive InfoNCE objective. Extensive experiments on four public benchmarks-RFMiD, ODIR, APTOS 2019, and IDRiD-demonstrate that MedRecord-CLIP yields statistically significant improvements over state-of-the-art baselines like RETFound and FLAIR on the multi-disease benchmarks (RFMiD and ODIR), and also leads on the IDRiD grading benchmark, while performing on par with the strongest baseline on APTOS 2019. Our approach highlights the critical value of integrating personalized clinical context to enhance the generalizability and interpretability of fundus foundation models.
