Related Experiment Video
Updated: Jan 9, 2026

13:19
Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
9.9K
Adaptive Pseudo Text Augmentation for Noise-Robust Text-to-Image Person Re-Identification
Lian Xiong1,2, Wangdong Li2, Huaixin Chen1
1School of Resources and Environment, University of Electronic Science and Technology of China, Chengdu 611731, China.
Sensors (Basel, Switzerland)
|December 11, 2025
Summary
This study introduces a novel text-to-image person re-identification (T2I-ReID) method that effectively handles noisy image-text pairs. By identifying and correcting misaligned data using pseudo-text generation, it improves cross-modal alignment for better pedestrian retrieval.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Text-to-image person re-identification (T2I-ReID) relies on aligned image-text pairs.
- Real-world data often contains noisy pairs due to coarse annotations and errors.
- Existing methods struggle with these data imperfections, limiting T2I-ReID performance.
Purpose of the Study:
- To develop a robust T2I-ReID method capable of handling noisy image-text pairs.
- To improve the accuracy and reliability of pedestrian retrieval from images/videos using textual descriptions.
- To enhance cross-modal alignment in the presence of data corruption.
Main Methods:
- Feature extraction using Contrastive Language-Image Pre-Training (CLIP).
- Token fusion model for fine-grained representations (Token Fusion Embedding - TFE).
- Gaussian Mixture Model (GMM) for noisy pair identification.
- Multimodal Large Language Model (MLLM) for pseudo-text generation to correct noisy data.
Main Results:
- The proposed method effectively identifies and mitigates the impact of noisy image-text pairs.
- Pseudo-text generation leads to more reliable visual-semantic associations.
- Significant improvements in T2I-ReID performance demonstrated across multiple benchmark datasets (CUHK-PEDES, ICFG-PEDES, RSTPReid).
- The model shows good compatibility with existing baseline methods.
Conclusions:
- The developed T2I-ReID approach successfully addresses the challenge of noisy image-text data.
- The noise identification and pseudo-text generation strategy enhances cross-modal alignment under noisy conditions.
- This work offers a promising direction for building more resilient T2I-ReID systems.
