Related Experiment Video
Updated: Oct 17, 2025

04:48
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
565
HybridCTrm: Bridging CNN and Transformer for Multimodal Brain Image Segmentation
Qixuan Sun1,2, Nianhua Fang1,2, Zhuo Liu3
1Key Laboratory for Ubiquitous Network and Service Software of Liaoning Province, Dalian, China.
Journal of Healthcare Engineering
|October 11, 2021
Summary
This study introduces HybridCTrm, a novel hybrid deep learning network combining Transformers and Convolutional Neural Networks (CNNs) for multimodal medical image segmentation. HybridCTrm demonstrates superior performance over traditional CNN-based methods, enhancing segmentation accuracy.
Area of Science:
- Medical Imaging
- Computer Vision
- Artificial Intelligence
Background:
- Multimodal medical image segmentation faces challenges with traditional deep learning methods like Convolutional Neural Networks (CNNs) due to limited long-range dependency modeling and generalization.
- Transformers have shown promise in image processing for improved generalization and performance, but CNNs offer advantages in rapid convergence and local feature representation.
Purpose of the Study:
- To propose a novel hybrid deep learning architecture, HybridCTrm, that integrates the strengths of both Transformers and CNNs for multimodal medical image segmentation.
- To evaluate the performance of HybridCTrm against existing methods on benchmark datasets and analyze the impact of Transformer depth on segmentation outcomes.
Main Methods:
- Developed HybridCTrm, a novel network architecture merging Transformer and CNN components for multimodal medical image segmentation.
- Conducted experiments on two benchmark datasets, comparing HybridCTrm against HyperDenseNet, a fully CNN-based network.
- Performed ablation studies to analyze the influence of Transformer depth on the network's performance and visualized segmentation results.
Main Results:
- HybridCTrm achieved superior performance compared to HyperDenseNet across most evaluation metrics on benchmark datasets.
- The study identified an optimal Transformer depth that balances performance and computational efficiency.
- Visualizations confirmed that the hybrid approach effectively improves segmentation quality by leveraging both global and local image features.
Conclusions:
- The proposed HybridCTrm network offers a significant advancement in multimodal medical image segmentation by effectively combining Transformer and CNN capabilities.
- HybridCTrm demonstrates improved generalization and performance, addressing limitations of purely CNN-based approaches.
- Further research into hybrid architectures holds promise for enhancing various medical image analysis tasks.

