Related Experiment Video
Updated: May 9, 2026

16:14
Trajectory Data Analyses for Pedestrian Space-time Activity Study
Published on: February 25, 2013
13.5K
Pedestrian Re-Identification Based on Fine-Grained Feature Learning and Fusion.
1Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen 518055, China.
Sensors (Basel, Switzerland)
|December 17, 2024
Summary
This study introduces a multimodal token-learning and alignment model (MTLA) for improved video-based pedestrian re-identification (Re-ID). The MTLA effectively fuses fine-grained features from visual and gait data, enhancing accuracy in cross-camera person identification.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Video-based pedestrian re-identification (Re-ID) faces challenges in learning effective representations due to video complexities like occlusion and blur.
- Existing multimodal fusion methods often operate at a global level, limiting their ability to capture fine-grained, complementary information.
Purpose of the Study:
- To propose a novel multimodal token-learning and alignment model (MTLA) for enhanced video-based pedestrian Re-ID.
- To achieve more effective pedestrian representation by learning, aligning, and fusing fine-grained features from multiple modalities.
Main Methods:
- The MTLA model incorporates a multimodal feature encoder to extract visual and gait features, learning and denoising fine-grained tokens.
- A token-based cross-modal alignment module aligns these multimodal features at the token level for capturing detailed semantic information.
- A correlation-aware fusion module integrates multimodal token features by learning inter- and intra-modal correlations for a unified representation.
Main Results:
- Extensive experiments on three benchmark datasets demonstrate the effectiveness of the proposed fine-grained feature alignment and fusion.
- The MTLA model achieved improvements exceeding 0.4 percentage points in mAP and Rank-K metrics compared to state-of-the-art approaches.
Conclusions:
- The proposed MTLA model significantly enhances video-based pedestrian Re-ID by leveraging fine-grained multimodal feature fusion.
- Fine-grained token-level alignment and correlation-aware fusion are crucial for capturing rich semantic information and improving Re-ID performance.

