Related Experiment Videos
Generalizable multi-modal medical image segmentation model using Multi-head Gated Cross Attention fusion
Chinnamgari Neeraja1, G Umamaheswara Reddy2, Gowri Thumbur3
1Research Scholar, Department of ECE, Sri Venkateswara University College of Engineering, Sri Venkateswara University, Tirupati, Andhra Pradesh 517502, India.
Abstract:
Automated medical image segmentation is an essential component of modern clinical diagnosis, enabling the precise delineation of anatomical features and abnormal tissue from various medical imaging techniques, including Magnetic Resonance Imaging (MRI) and Computed Tomography (CT) scans. However, achieving reliable segmentation across multi-modal medical images remains difficult because of variations in imaging properties, feature distributions, and modality-specific information. Existing segmentation methods often exhibit limited generalization and insufficient cross-modal feature integration, resulting in reducing segmentation accuracy. To resolve these concerns, the proposed framework employs an advanced attention mechanism-driven deep learning network to effectively learn complex heterogeneous features in medical images. The process begins by collecting the required multi-modal images from standard public databases, which are then directly provided to the segmentation model for performing multi-modal segmentation. The segmentation process is executed using Multi-head Gated Cross Attention Fusion Encoder-based Adaptive Transformer Mobile-UNet++ with Consistency Loss Function (MGAFE-ATMU++-CLF) for accurate segmentation of multi-modal medical images. The proposed model incorporates a multi-head gated cross-attention fusion mechanism throughout the encoder to efficiently extract complementary information from different imaging modalities while preserving discriminative features. Furthermore, a Renovated Puma Optimizer (RPO) is designed to fine-tune the critical network hyperparameters for the MGAFE-ATMU++ framework, thereby enhancing the segmentation capability of the developed approach. The evaluation results indicate the introduced MGAFE-ATMU++-CLF framework delivers improved segmentation accuracy by effectively extracting heterogeneous features, integrating cross-modal information, and accurately delineating anatomical boundaries, highlighting its potential for reliable analysis of multi-modal medical images.