Related Experiment Video
Updated: May 21, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
MALFM-Captioner: A Multipath Alignment Learning for Image Captioning With Feature Mask
Abstract:
Diffusion-based image captioning models effectively mitigate the token dependency issue inherent in autoregressive methods. However, the noise introduced in diffusion methods weakens sentence information, resulting in insufficient ability of image-text feature alignment. To address this issue, we propose a multipath alignment learning for image captioning with feature mask (MALFM-Captioner) method. Leveraging both global and regional visual features, we first introduce a feature masked module (FMM) that enables the model to reconstruct masked visual information during training, thereby enhancing its capability to learn discriminative visual representations. Concurrently, we perform cross-attention between deep image and text features and fuse the outputs via weighted summation, effectively mitigating the image-text misalignment problem inherent to single-path paradigms. Furthermore, we design a concise yet effective gated feature fusion module (GFFM) to integrate complementary feature alignment results, improving the accuracy and semantic fidelity of generated captions. Extensive experiments on the MS COCO and Flickr 30K datasets demonstrate that MALFM-Captioner achieves 0.9% and 1.9% improvements in Bleu-4 and CIDEr metrics, respectively, and exhibits competitive performance against state-of-the-art models (e.g., DDCap and Bit Diffusion).