Related Experiment Videos
Surg-NAT+: negation-aware vision-language refinement for fine-grained surgical understanding
Kun Yuan1, Yutong Cao2, Yunxi Tang3
1CAMMA, University of Strasbourg, Strasbourg, France. kun.yuan@ext.ihu-strasbourg.eu.
Summary
Surg-NAT+ enhances surgical vision-language models by improving their ability to understand negation in text prompts, leading to better few-shot tool recognition. This framework boosts performance in fine-grained surgical tasks.
Area of Science:
- Computer Vision
- Machine Learning
- Medical AI
Background:
- Surgical vision-language foundation models leverage large datasets for generalizable representations.
- Current models struggle with semantic limitations, particularly distinguishing positive and negative text assertions.
- This hinders fine-grained surgical tasks requiring precise discrimination of localized features.
Purpose of the Study:
- To introduce Surg-NAT+, a few-shot vision-language adaptation framework.
- To enhance foundation models for fine-grained surgical tasks by addressing negation limitations.
- To improve the reliability and performance of surgical vision-language models.
Main Methods:
- Developed a negation-aware contrastive objective to differentiate affirmative and negated prompts.
- Utilized multi-level adapter fusion (MAF) for hierarchical semantic refinement in text encoders.
- Implemented a fine-grained self-distillation objective for improved visual grounding via consistency between global and local representations.
Main Results:
- Surg-NAT+ achieved state-of-the-art performance on the Cholec80 dataset for few-shot surgical tool recognition.
- The framework consistently outperformed existing baselines across all few-shot regimes.
- Qualitative analyses confirmed more accurate visual grounding and enhanced semantic separability.
Conclusions:
- Surg-NAT+ is a lightweight, negation-aware framework improving fine-grained recognition in surgical vision-language models.
- It offers an efficient and semantically robust adaptation pathway for surgical domains.
- The framework excels in few-shot, multi-label tool recognition settings.
Related Concept Videos
Vision
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
Prosopagnosia
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...