Related Experiment Videos
Visible-Infrared Fusion Based on CNN and Deformable Transformer
Xiaoyi Wang1,2, Xiansong Gu1, Bin Li2
1College of Opto-Electronic Engineering, Changchun University of Science and Technology, Changchun 130022, China.
None:
To address the limitations of traditional methods in feature extraction and multi-modal information fusion, this paper proposes an infrared-visible image object detection architecture that integrates Convolutional Neural Networks (CNNs) and Deformable Transformers. This method leverages the advantages of CNN in local feature modeling and the capabilities of Transformer in capturing global contextual information, facilitating the fusion of semantic consistency and structural details across modalities. By introducing a detection-aware multi-task optimization mechanism, the model improves object detection in challenging scenarios such as low-light conditions, occlusion, and complex backgrounds. Experiments on multiple standard datasets, including M3FD and LLVIP, indicate that the proposed method achieves competitive or better performance than the compared methods in key metrics such as mAP. Specifically, our method obtains the best mAP50 among the evaluated methods with an mAP50 of 74.2% on the M3FD dataset and 98.6% on the LLVIP dataset, surpassing the second-best PIAFusion by 4.3% and 2.5% respectively. These quantitative results support the practicality and effectiveness of our approach in the evaluated complex environments.