Related Experiment Video
Updated: May 19, 2026

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
Application of transformer models in medical image segmentation: a narrative review
Yuan Xu1, Yang Peng1, Chi Zhang2
1School of Computer Science, Sichuan Normal University, Chengdu, China.
Background And Objective:
Transformer-based medical image segmentation plays a crucial role in healthcare applications, facilitating precise diagnosis, treatment planning, and disease monitoring. Traditionally, convolutional neural networks (CNNs), which excel at local feature extraction, have dominated this field. However, they have a limited ability to capture the long-range dependencies within images and thus face difficulty in handling the complex, interconnected structures present in medical data. Transformer modeling, as an advanced tool in natural language processing (NLP), has also demonstrated its value in computer vision tasks. With its increasing popularity, research on its application in medical imaging has grown significantly. Nowadays, several models have demonstrated that combining CNNs and transformers can effectively capture both local and global information, thereby enhancing segmentation performance. A review was conducted to characterize the research on transformer models applied in medical image segmentation.
Methods:
Databases including Google Scholar, arXiv, ResearchGate, Microsoft Academic, PubMed, and Semantic Scholar, as well as large language models including ChatGPT and DeepSeek, were used to search for the latest developments in this field. Specifically, English language literature published from 2021 to 2025 was included in the review.
Key Content And Findings:
In this investigation, we explored a variety of methods for integrating Transformer with traditional U-shaped architectures, as well as the efficiency disparities exhibited by different approaches. Through this investigation, we conducted research centered on the U-shaped architectures of pure Transformer and hybrid Transformer, and analyzed models that have been proven to possess outstanding performance. Segmentation models can be divided into pure transformer and hybrid architectures. In this review, the applications of transformer models in medical segmentation were examined, the performance of these models on different datasets was summarized, and advanced strategies were quantitatively analyzed. Multimodal large language models have been reported in the recent literature, signifying that these no longer constitute a speculative technology but rather an emerging development in the development field.
Conclusions:
The number of pure transformer models is relatively small compared with that of and hybrid models, with the latter generally demonstrating superior performance. Among hybrid models, those integrating the transformer into the decoder typically outperform other hybrid variants; however, the number of such models is also limited. Public datasets in medical image segmentation are highly diverse, with corresponding datasets available for a variety of data modalities. Further research will focus on lightweight structural design, prior knowledge integration, data utilization efficiency enhancement, and foundational model scale expansion to improve model practicality and generalization in low-compute, low-data scenarios. Finally, the future developments for improving the effectiveness of the transformer models in the medical field were summarized.