Related Experiment Video
Updated: Apr 29, 2026

07:11
Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping
Published on: December 8, 2023
2.6K
SAVLT: Structure-Aware Vision-Language Tuning for Multi-Center Cervical OCT Diagnosis
IEEE Journal of Biomedical and Health Informatics
|April 27, 2026
Summary
We developed SAVLT, a novel framework for cervical optical coherence tomography (OCT) analysis. This structure-aware tuning method improves diagnostic accuracy by focusing on tissue integrity, overcoming challenges posed by artifacts in vision-language models.
Area of Science:
- Biomedical Imaging
- Artificial Intelligence in Medicine
- Medical Diagnostics
Background:
- Cervical optical coherence tomography (OCT) provides high-resolution tissue visualization but faces diagnostic challenges with limited supervision.
- Vision-language models (VLMs) show promise but struggle with artifacts that mimic biological structures, hindering pathological analysis.
- Artifacts in OCT images can distract models, obscuring crucial details of layer degradation essential for accurate diagnosis.
Purpose of the Study:
- To introduce SAVLT, a structure-aware tuning framework for adapting VLMs in cervical OCT analysis.
- To enhance diagnostic reliability by addressing confounding artifacts and improving focus on tissue pathology.
- To enable robust few-shot generalization and clinical interpretability for foundation models in heterogeneous OCT imaging.
Main Methods:
- SAVLT employs parameter-efficient fine-tuning to adapt VLMs.
- A region-aware spatial attention (RaSA) module is introduced to enforce spatial constraints and purify visual representations.
- A dual-constraint objective combines image-text alignment with visual prototypes for stable optimization across domains.
Main Results:
- SAVLT effectively shifts VLM attention from global matching to anatomical grounding.
- The RaSA module successfully removes non-biological noise, restoring focus on intra-tissue structural integrity.
- The framework demonstrated robust few-shot generalization and clinical interpretability across multi-center datasets.
Conclusions:
- SAVLT establishes a reliable paradigm for deploying foundation models in cervical OCT imaging.
- The proposed structure-aware tuning framework significantly improves diagnostic accuracy by mitigating artifact interference.
- SAVLT offers a promising solution for trustworthy AI-driven diagnostics in medical imaging with limited supervision.

