Related Experiment Video
Updated: Jun 29, 2025

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
7.6K
AA-RGTCN: reciprocal global temporal convolution network with adaptive alignment for video-based person
Yanjun Zhang1, Yanru Lin2, Xu Yang3
1School of Cyberspace Science and Technology, Beijing Institute of Technology, Beijing, China.
Frontiers in Neuroscience
|April 9, 2024
Summary
This study introduces a new method for video-based person re-identification (Re-ID) that improves accuracy by addressing frame misalignment and enhancing temporal feature modeling. The proposed Adaptive Alignment-Reciprocal Global Temporal Convolution Network (AA-RGTCN) achieves state-of-the-art results on benchmark datasets.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Person re-identification (Re-ID) is crucial for retrieving individuals across different camera views.
- Video-based Re-ID leverages both spatial and temporal features, but existing methods suffer from redundant spatial descriptions and insufficient temporal modeling.
- Current approaches often ignore frame misalignment and rely on fixed-length sequences, limiting temporal representation.
Purpose of the Study:
- To develop a novel network architecture for video-based person re-identification.
- To address the challenges of frame misalignment and insufficient temporal feature modeling in existing Re-ID methods.
- To improve the accuracy and robustness of person re-identification in video surveillance.
Main Methods:
- Proposed the Adaptive Alignment-Reciprocal Global Temporal Convolution Network (AA-RGTCN).
- Introduced an Adaptive Alignment block to dynamically adjust frame positions for optimal temporal modeling.
- Developed a Reciprocal Global Temporal Convolution Network to capture robust temporal features across varying time intervals and directions.
Main Results:
- Achieved 85.9% mAP and 91.0% Rank-1 accuracy on the MARS dataset.
- Obtained 90.6% Rank-1 accuracy on the iLIDS-VID dataset.
- Reached 96.6% Rank-1 accuracy on the PRID-2011 dataset, outperforming state-of-the-art methods.
Conclusions:
- The AA-RGTCN effectively handles frame misalignment and models discriminative temporal representations.
- The proposed method demonstrates superior performance in video-based person re-identification compared to existing approaches.
- The Reciprocal Global Temporal Convolution Network contributes to robust temporal feature extraction for Re-ID tasks.

