Related Experiment Videos
Concept-Enhanced Multi-Scale Cross-Modal Alignment for Medical Visual Representation Learning
Abstract:
Medical Vision-Language Pre-training (Med VLP) on paired medical images and reports has emerged as a promising direction for learning visual representations. However, current alignment approaches remain insufficient for learning fine-grained pathological details, largely due to the inherent difficulty of tokenizing reports without losing accurate and complete radiological concepts. To ad dress this, we propose a novel Concept-Enhanced Medical Vision-Language Alignment (CoMA) framework, which facilitates the acquisition of medical concepts (e.g., pathological findings) by effectively leveraging the radiologists' experience embedded in the text. Specifically, based on radiologists' cognitive framework, we design a Concept Clause Decomposition method to extract semantically complete descriptions of pathological findings or radiology manifestations from medical reports as medical concept clauses. These clauses are then utilized within a multi-granularity cross-modal alignment framework to enhance medical concept perception and fine-grained medical vision-language alignment. Beyond Global Instance-level Alignment (GIA), we introduce Concept-level Token-wise Alignment (CTA) via a bidirectional cross-attention mechanism, which enhances semantic correspondence at the pathological region level by aligning image patches with concept clauses. At a higher semantic level, we establish a bi-level proto type alignment mechanism, comprising Instance-level Prototype Alignment (IPA) and Concept-level Prototype Alignment(CPA).Through cross-modal prototype clustering, this mechanism pulls together unpaired high-level global features and low-level concept features across image and text, respectively. Extensive experiments across seven super vised downstream dataset-task settings demonstrate that CoMA achieves consistent and competitive performance, particularly in limited-annotation settings.