Related Experiment Videos
PatchCLIP enables region specific contrastive health record and image joint training with patch embedding loss
Sheethal Bhat1,2, Awais Mansoor3, Bogdan Georgescu3
1Pattern Recognition Lab, Friedrich-Alexander-Universität, Erlangen-Nürnberg, 91058, Erlangen, Germany. sheethal.bhat@fau.de.
Abstract:
Vision-Language (VL) models such as Contrastive Language-Image pretraining (CLIP) have shown remarkable zero-shot classification capabilities by jointly learning from large-scale image-text datasets using multimodal self-supervised learning (SSL). However, while these models capture strong global semantics, they often struggle with fine-grained spatial understanding, thereby limiting their effectiveness in downstream tasks like object detection and medical abnormality localization2. To address this limitation, we propose Patch-CLIP, a novel VL framework that introduces a contrastive loss aligning image patch-level embeddings with text embeddings. Unlike conventional VL approaches that only leverage global image representations, our method utilizes local patch-level features to encode spatial context, enabling effective learning of localization cues. Applied to two Chest X-ray (CXR) datasets, Patch-CLIP achieves state-of-the-art (SOTA) performance across eight abnormality detection tasks. Furthermore, the resulting patch prediction maps substantially reduce false positives at comparable sensitivity levels compared to standard saliency-based methods, providing more precise and interpretable localization of key findings. The code is available at https://github.com/Siemens-Healthineers/patch-clip.
Related Concept Videos
Long-patch Base Excision Repair
Patch Clamp
In this method, a glass micropipette containing electrolyte solution is tightly sealed against a small portion of the cell membrane. As a result, a patch of the cell...
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...
Improving Translational Accuracy
Improving Translational Accuracy
Clipper Circuit
The operation of a clipper circuit can be exemplified by analyzing a dual-clipper configuration setup that integrates two ideal diodes, each paired with a biasing...