Related Experiment Videos
HiVLR: Hierarchical Vision-Language Reasoning for interpretable zero-shot radiography image understanding
Xilin Dang1, Kang Li2, Pheng Ann Heng1
1Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong.
None:
Medical vision-language pre-training on image-report pairs has shown great potential to facilitate downstream image understanding tasks. However, prior approaches commonly exhibited limited accuracy on zero-shot image tasks, and lacked sufficient interpretability for unseen disease diagnosis, posing substantial usability concerns and trust issues for safety-critical medical applications. To alleviate them, we revisit how human doctors reason from a patient's radiology image for diagnosis, and propose a Hierarchical Vision-Language Reasoning (HiVLR) framework based on the clinical diagnostic workflow. In specific, we structure feature investigation into sequential rounds of thinking, i.e., (1) spotting the suspicious pathology observations (e.g., obscure) from all visual and textual inputs first and then (2) determining possible diagnostic findings (e.g., pneumonia) that match all pathology observations, to derive accurate disease predictions without compromising transparency in model decision making. Each round of thinking needs to analyze the inputted visual and textual embeddings by coarsely aligning them with cross-attention to establish global correspondences, highlight region-level visual features containing specific clinical content by prompt tuning-enabled fine-grained filtering, and then interpret the visual features in a condensed understanding to derive diagnostic-pertinent discoveries. Importantly, we enforce concept-level cross-modal compliance by ensuring that visual and textual features corresponding to the same clinical content are semantically consistent across concept dimensions (e.g., texture, shape, border). Based on this, we attach a concept-based interpretable diagnosis block to improve the accuracy and interpretability in downstream tasks simultaneously. Experiments showed that our approach greatly outperformed competing approaches on diverse zero-shot image tasks with superior interpretability.
Related Concept Videos
Vision
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
Imaging Studies III: Computed Tomography
X-ray Imaging
Depth Perception and Spatial Vision