Related Experiment Video
Updated: Aug 7, 2025

04:48
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
471
Transformer-based multi-task learning for classification and segmentation of gastrointestinal tract endoscopic images
Suigu Tang1, Xiaoyuan Yu1, Chak Fong Cheang1
1Faculty of Innovation Engineering-School of Computer Science and Engineering, Macau University of Science and Technology, Macao Special Administrative Region of China.
Computers in Biology and Medicine
|March 12, 2023
Summary
A novel Multi-task Network (TransMT-Net) improves gastrointestinal (GI) disease diagnosis by combining CNN and transformer features for accurate classification and segmentation. Active learning significantly enhances performance, even with limited labeled endoscopic images.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Gastroenterology
Background:
- Convolutional Neural Networks (CNNs) are widely used for gastrointestinal (GI) tract disease identification in endoscopic images.
- CNNs struggle with ambiguous lesion classification and require large labeled datasets for effective training.
- Existing models face limitations in accurately distinguishing similar lesions and generalizing with insufficient data.
Purpose of the Study:
- To develop an advanced deep learning model for improved accuracy in classifying and segmenting GI tract lesions.
- To address the challenge of limited labeled data in training diagnostic models for endoscopic imaging.
- To enhance the diagnostic capabilities for GI diseases using a novel network architecture and active learning.
Main Methods:
- Proposed a Multi-task Network (TransMT-Net) integrating CNNs for local features and transformers for global features.
- Implemented simultaneous learning for both lesion classification and segmentation tasks.
- Utilized active learning to efficiently manage and leverage limited labeled datasets.
Main Results:
- TransMT-Net achieved 96.94% accuracy in classification and 77.76% Dice Similarity Coefficient in segmentation.
- The model demonstrated superior performance compared to other existing models on the test dataset.
- Active learning enabled comparable performance with only 30% of the training data.
Conclusions:
- TransMT-Net shows significant potential for accurate GI tract endoscopic image analysis.
- The integration of active learning effectively mitigates the need for extensive labeled datasets.
- The proposed approach enhances diagnostic accuracy and efficiency in identifying GI diseases.

