Related Experiment Video
Updated: Jan 19, 2026

Lung CT Segmentation to Identify Consolidations and Ground Glass Areas for Quantitative Assesment of SARS-CoV Pneumonia
Published on: December 19, 2020
TA-MedSAM: Text-augmented improved MedSAM for pulmonary lesion segmentation
Siyuan Tang1, Siriguleng Wang2, Gang Xiang3
1College of Mathematical Sciences, Inner Mongolia Normal University, Hohhot, Inner Mongolia 010022, China; College of Computer Science and Technology, Baotou Medical College, Baotou, Inner Mongolia 014040, China.
Abstract:
Accurate segmentation of lung lesions is critical for clinical diagnosis. Traditional methods rely solely on unimodal visual data, which limits the performance of existing medical image segmentation models. This paper introduces a novel approach, Text-Augmented Medical Segment Anything Module(TA-MedSAM), which enhances cross-modal representation capabilities through a vision-language fusion paradigm. This method significantly improves segmentation accuracy for pulmonary lesions with challenging characteristics including low contrast, blurred boundaries, complex morphology, and small size. Firstly, we introduce a lightweight Medical Segment Anything Model (MedSAM) image encoder and a pre-trained ClinicalBERT text encoder to extract visual and textual features, This design preserves segmentation performance while reducing model parameters and computational costs, thereby enhancing inference speed. Secondly, a Reconstruction Text Module is proposed to focus the model on lesion-centric textual cues, strengthening semantic guidance for segmentation. Thirdly, we develop an effective Multimodal Feature Fusion Module that integrates visual and textual features using attention mechanisms, and introduce a feature alignment coordination mechanism to mutually enhance heterogeneous information across modalities, and a Dynamic Perception Learning Mechanism is proposed to quantitatively evaluate fusion effectiveness, enabling optimal fused feature selection for improved segmentation accuracy. Finally, a Multi-scale Feature Fusion Module combined with a Multi-task Loss Function enhances segmentation performance for complex regions. Comparative experiments demonstrate that TA-MedSAM outperforms state-of-the-art unimodal and multimodal methods on QaTa-COV19, MosMedData+ , and private dataset. Extensive ablation studies validate the efficacy of our proposed components and optimal hyperparameter combinations.

