Related Experiment Video
Updated: May 2, 2026

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
Cross-modal transformer fusion via local sampling for drone RGB-infrared object detection
Herong Qi1, Xuanyu Xiang1, Hui Qin1
1School of Artificial Intelligence and Automation, National Key Laboratory of Multispectral Information Intelligent Processing Technology, Huazhong University of Science and Technology, Wuhan, 430074, China.
This study introduces a new network for drone-based object detection, improving feature fusion from high-resolution images. The method enhances the detection of small or complex objects by preserving local details.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Remote Sensing
Background:
- Visible and infrared feature fusion is crucial for drone-based RGB-IR object detection.
- Existing transformer-based methods often use downsampled features, losing critical local details and reducing accuracy, especially for small or complex objects.
Purpose of the Study:
- To propose a novel Cross-Modal Transformer Fusion via Local Sampling (CTFLS) network for improved drone-based object detection.
- To address the limitations of low-resolution feature fusion in current transformer models.
Main Methods:
- A two-stream strategy is employed to capture intra-modal and inter-modal information from high-resolution feature maps.
- A Local Cascade Transformer (LCT) module with local intra-modal and cross-modal transformer blocks extracts rich information from high-resolution features.
- A Detail-Enhanced Mixed-Convolution Attention (DMA) module enhances fused feature representation for subtle textures and global context.
Main Results:
- The proposed CTFLS network effectively captures both intra-modal and inter-modal information from high-resolution feature maps.
- Local fusion via cross-attention on sampled local features explores deeper complementary relationships between modalities.
- The DMA module improves the representation of fused features, capturing subtle textures and global context.
Conclusions:
- The CTFLS network outperforms state-of-the-art methods in drone-based object detection.
- The method demonstrates significant improvements in detecting small objects and distinguishing objects with similar structures.
Related Concept Videos
Light Acquisition
Transformation
Transformers with Off-Nominal Turns Ratios
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
