Related Experiment Video
Updated: May 9, 2025

07:12
Profiling Maternal Behavior Responses During Whole-Brain Imaging
Published on: January 24, 2025
559
Channel-shuffled transformers for cross-modality person re-identification in video
Rangwan Kasantikul1,2, Worapan Kusakunniran3, Qiang Wu4
1Faculty of Information and Communication Technology, Mahidol University, 999 Phuttamonthon 4 Road, Salaya, 73170, Nakhon Pathom, Thailand.
Scientific Reports
|April 29, 2025
Summary
This study introduces the Hybrid Channel-Shuffled Transformer Net (HCSTNET) for person re-identification (Re-ID) across different visual modalities. The novel approach effectively reduces overfitting and improves temporal feature extraction in surveillance applications.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Person re-identification (Re-ID) across different modalities (e.g., daylight vs. night-vision) is critical for surveillance.
- Multi-frame information is essential for robust Re-ID, as single frames are often unreliable.
- Standard transformers face scaling challenges due to high channel requirements for feature encoding, leading to overfitting and instability.
Purpose of the Study:
- To propose a novel Channel-Shuffled Temporal Transformer (CSTT) integrated with a ResNet backbone, forming the Hybrid Channel-Shuffled Transformer Net (HCSTNET).
- To address the scaling challenges and overfitting issues in transformer-based Re-ID models.
- To improve temporal information extraction and attention mechanisms for multi-frame person re-identification.
Main Methods:
- Developed the Hybrid Channel-Shuffled Transformer Net (HCSTNET) by combining a ResNet backbone with a novel Channel-Shuffled Temporal Transformer (CSTT).
- Replaced fully connected layers in standard multi-head attention with ShuffleNet-like structures for ResNet integration.
- Utilized channel-grouping and channel-shuffling within ShuffleNet-like structures to reduce parameters and enhance attention learning.
Main Results:
- The HCSTNET, particularly the temporal transformer with channel-shuffling, demonstrated measurable improvement over baseline methods (simple frame averaging) on the SYSU-MM01 dataset.
- Channel-shuffling and parameter reduction through ShuffleNet-like structures effectively mitigated overfitting and improved attention.
- Investigated optimal partitioning strategies for feature maps within the proposed architecture.
Conclusions:
- The proposed HCSTNET effectively handles multi-frame person re-identification across different modalities.
- Channel-shuffling integrated with transformer attention provides a viable solution to scaling challenges and overfitting in Re-ID.
- The method offers a significant advancement for surveillance applications requiring robust person re-identification.

