Related Experiment Video
Updated: Jun 21, 2025

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
Published on: October 27, 2023
Reliable object tracking by multimodal hybrid feature extraction and transformer-based fusion.
Hongze Sun1, Rui Liu1, Wuque Cai1
1Clinical Hospital of Chengdu Brain Science Institute, MOE Key Lab for NeuroInformation, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu 611731, China.
This study introduces a novel multimodal hybrid tracker (MMHT) for reliable visual object tracking. The MMHT effectively addresses challenges like low light and clutter by integrating frame-event data for improved feature modeling.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Robotics
Background:
- Visual object tracking using visible light faces challenges like low light, high dynamic ranges, and background clutter.
- Existing multimodal approaches often use local feature interactions, limiting feature modeling potential.
- Reliable object tracking requires robust feature extraction and fusion across diverse visual cues.
Purpose of the Study:
- To propose a novel multimodal hybrid tracker (MMHT) for robust single object tracking.
- To enhance feature modeling by integrating frame-event data and advanced neural networks.
- To improve tracking performance in complex visual environments.
Main Methods:
- Utilized a hybrid backbone combining an artificial neural network (ANN) and a spiking neural network (SNN) for feature extraction.
- Employed a unified encoder to align features from different visual modalities.
- Developed an enhanced transformer-based module with attention mechanisms for multimodal feature fusion.
Main Results:
- The MMHT model effectively constructs a multiscale and multidimensional visual feature space.
- Achieved discriminative feature modeling, leading to enhanced tracking accuracy.
- Demonstrated competitive performance against state-of-the-art methods in extensive experiments.
Conclusions:
- The proposed MMHT model effectively addresses challenges in visual object tracking.
- Integrating frame-event data with hybrid ANN-SNN architectures improves tracking reliability.
- The study highlights the potential of multimodal fusion for robust computer vision tasks.
Related Concept Videos
Extraction: Advanced Methods
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...

