Related Experiment Video
Updated: Mar 15, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Revisiting Text-Based Person Retrieval: Mitigating Annotation-Induced Mismatches with Multimodal Large Language
Zihang Han1, Chao Zhu1, Mengyin Liu1
1School of Computer and Communication Engineering, University of Science and Technology Beijing, Beijing 100083, China.
Existing text-based person retrieval benchmarks contain ambiguous descriptions, leading to inaccurate model evaluations. This study introduces an annotation refinement framework using multimodal large language models to generate distinctive descriptions, improving benchmark quality and TBPR model performance.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Text-based person retrieval (TBPR) relies on high-quality benchmarks for accurate model evaluation.
- Current TBPR benchmarks suffer from ambiguous annotations, where similar images have identical descriptions, causing evaluation errors.
- Individual image annotation without considering similar images hinders the creation of distinctive descriptions.
Purpose of the Study:
- To propose an annotation refinement framework to enhance the quality of TBPR benchmarks.
- To mitigate annotation-induced mismatches in TBPR model evaluation.
- To improve the discriminative power of textual descriptions for person images.
Main Methods:
- Automated identification of image sets prone to mismatches using TBPR models.
- Leveraging multimodal large language models (MLLMs) to process multiple images simultaneously.
- Generating distinctive descriptions for each image and replacing original annotations to improve quality.
Main Results:
- The proposed framework effectively improves annotation quality in TBPR benchmarks.
- Experiments on CUHK-PEDES, RSTPReid, and ICFG-PEDES validate the method's effectiveness.
- More discriminative captions generated by the framework benefit mainstream TBPR models.
Conclusions:
- The annotation refinement framework significantly enhances TBPR benchmark quality.
- Improved benchmarks lead to more accurate evaluation of TBPR models.
- The enhanced benchmark datasets will be publicly released to benefit the research community.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Mismatch Repair
Mismatch Repair
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...