Related Experiment Video
Updated: Jan 15, 2026

04:48
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
732
Learning Hierarchically Consistent Disentanglement with Multi-Channel Augmentation for Public Security-Oriented
Yu Ye1,2, Zhihong Sun3, Jun Chen1,2
1National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University, Wuhan 430072, China.
Sensors (Basel, Switzerland)
|October 16, 2025
Summary
This study introduces a new method for sketch re-identification (Re-ID) to match sketch images with photos. The approach effectively bridges the modality gap, improving accuracy in public security applications.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Sketch re-identification (Re-ID) is vital for public security tasks like criminal investigations and missing person searches.
- A significant challenge in Re-ID is the modality gap between sketch images and real photographs, especially due to color information differences.
- Existing methods struggle to extract robust, modality-invariant features essential for accurate cross-modal matching.
Purpose of the Study:
- To propose a novel network architecture for sketch re-identification (Re-ID) that overcomes the modality gap.
- To mitigate the interference of color bias in cross-modal matching between sketches and photographs.
- To develop a method for decomposing pedestrian representations into modality-invariant and modality-specific features.
Main Methods:
- A novel network architecture integrating multi-channel augmentation and hierarchically consistent disentanglement learning.
- A multi-channel augmentation module to reduce color bias interference.
- A modality-disentangled prototype (MDP) module for feature-level decomposition and a cross-layer decoupling consistency constraint.
Main Results:
- The proposed approach significantly improves performance on sketch re-identification tasks.
- Experimental results on two public datasets demonstrate the superiority over state-of-the-art methods.
- The network effectively bridges the modality gap, extracting discriminative modality-invariant features.
Conclusions:
- The novel network architecture effectively addresses the challenges of sketch re-identification.
- The integration of multi-channel augmentation and disentanglement learning enhances cross-modal matching accuracy.
- The proposed method offers a promising solution for real-world applications in public security.
