Related Experiment Video
Updated: Aug 29, 2026

Positron Emission Tomography-based Dose Painting Radiation Therapy in a Glioblastoma Rat Model using the Small Animal Radiation Research Platform
Published on: March 24, 2022
MGTP-Seg: A Mask Guidance and Text Prompting Network for Gross Tumor Volume Segmentation in Esophageal Cancer
Chengwei Chen1, Hongfei Sun2, Yuxuan Yao1
1Department of Biomedical Engineering, Air Force Medical University, No. 169 Changle West Road, Xi'an, 710032, China.
Abstract:
Accurate delineation of the gross tumor volume (GTV) is critical for determining the efficacy of radiotherapy in esophageal cancer. Conventional segmentation methods either rely solely on end-to-end learning from imaging features, which often struggle to address small tumor volumes and ambiguous boundaries, or incorporate coarse masks as spatial priors but fail to integrate the pathological semantics essential for clinical decision-making. This disconnect can lead to segmentation results that are poorly aligned with clinical practice. To address these limitations, this study proposed the Mask Guidance and Text Prompting Segmentation (MGTP-Seg) framework. Based on UNETR, MGTP-Seg innovatively integrates a mask guidance branch and a text prompting branch. The mask guidance branch utilizes pre-segmented masks from nnU-Net to provide spatial priors, while the text prompting branch dynamically integrates clinically relevant semantics from large language models into the visual feature space via learnable prompt tuning. Through adaptive fusion and bidirectional alignment, these branches enable synergistic integration of imaging details, spatial priors, and high-level clinical knowledge in an end-to-end manner. Experiments on multi-center datasets confirm that MGTP-Seg delivers accurate and robust segmentation on both internal and external validation sets. This work demonstrates that MGTP-Seg not only provides an accurate, interpretable, and clinically relevant solution for automatic GTV delineation, but also offers a novel methodological framework to fuse spatial priors with semantic knowledge in medical image analysis.
