Related Experiment Videos
Cross-modality fusion with local and global semantic feature correlation for weakly distinguishable object detection
Maozhen Liu1, Xiaoguang Di1, Ximing Li1
1Control and Simulation Center, Harbin Institute of Technology, Harbin, Heilongjiang, 150080, China; National Key Laboratory of Modeling and Simulation for Complex Systems, Harbin, Heilongjiang, 150080, China.
Abstract:
To address the issue of missed detections caused by the similarity between weakly distinguishable objects and the background in both visible and infrared images, we utilize the correlation between local and global semantic features to find the difference clues between objects and the background. Specifically, we propose a novel Cross-Modality Fusion with Local and Global Semantic Feature Correlation Object Detection Network (CFLGNet), which effectively enhances the distinguishability. Firstly, we design a Semantic-Attention Mamba Fusion Block called SAMFB, which maps the cross-modal features into a hidden state space for interaction, and combines local and global semantic feature correlation to enhance the distinguishing features between the object and the background. SAMFB contains two branches: The Local Semantic Feature Correlation Module (LSCM) obtains an effective representation of local fine semantic feature through bilinear matrix computation, and this feature representation is embedded into another Global Semantic Feature Correlation Module with mamba (GSCM). GSCM captures the irregular spatial correlation through the scanning strategy of visual mamba, and then the local semantic feature extracted by the LSCM module are used as weights to suppress the spatially adjacent features with different representations. In addition, we propose a Lcal Feature Enhancement module based on Topological Invariants(LFETI) to reduce the problem of missed detection caused by incomplete object information, such as occlusion or truncation. Finally, the semantic feature correlation is combined with the original feature to enhance the distinguishing feature between the object and the background. Extensive experiments on the VEDAI and DroneVehicle benchmark datasets demonstrate that the proposed CFLGNet exhibits remarkable performance.
Related Concept Videos
Attenuated Total Reflectance (ATR) Infrared Spectroscopy: Overview
The ATR process begins by directing a beam...
Infrared (IR) Spectroscopy: Overview
Different compounds display unique properties due to their...
IR Frequency Region: Fingerprint Region
The...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Light Acquisition