Related Experiment Video
Updated: May 14, 2026

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
Text-Visible/Infrared Person Retrieval: Attribute-Guided Feature Decoupling and Collaborative Alignment and a Unified
Summary
This study introduces Text-Visible/Infrared person retrieval, enabling matching text descriptions to both visible and infrared images. The novel Attribute-guided feature decoupling and Collaborative Alignment Network (ACANet) achieves accurate cross-modal alignment for improved low-light person identification.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Current text-to-image person retrieval methods are limited to visible imagery, failing in low-light conditions.
- Infrared imaging is crucial for many visual systems, necessitating text matching across visible and infrared modalities.
- Heterogeneous visual characteristics of visible and infrared images pose challenges for unified cross-modal retrieval frameworks.
Purpose of the Study:
- To introduce a new task: Text-Visible/Infrared person retrieval.
- To propose a novel network, ACANet, for unified text-to-visible/infrared image matching.
- To establish a new benchmark dataset for advancing this research area.
Main Methods:
- Developed the Attribute-guided feature decoupling and Collaborative Alignment Network (ACANet) for unified cross-modal alignment.
- Decoupled visible image color and texture features, integrating them with infrared features for enhanced alignment.
- Extended masked language modeling to a cross-modal paradigm for fine-grained alignment across multiple image modalities.
Main Results:
- The proposed ACANet demonstrated superior performance compared to state-of-the-art methods on the new MM01LLCM-Text dataset.
- Successfully addressed the challenge of matching text descriptions to heterogeneous visible and infrared images.
- Validated the effectiveness of feature decoupling and cross-modal alignment strategies.
Conclusions:
- ACANet provides an effective unified framework for Text-Visible/Infrared person retrieval.
- The developed MM01LLCM-Text dataset facilitates further research in this domain.
- This work advances the capability of person retrieval in challenging low-light and multi-modal scenarios.
Related Concept Videos
IR Frequency Region: Fingerprint Region
IR spectra are divided into two main regions: the diagnostic region and the fingerprint region. The diagnostic region of the spectrum lies above 1500 cm−1. The absorptions resulting from single-bond vibrations of the N–H, C–H, and O–H stretch at higher wavenumbers and appear on the left side of the spectrum. The stretching absorptions of the C≡C and C≡N occur between 2100–2300 cm−1. In contrast, those arising from stretching absorptions of the C=O, C=N, and C=C occur between 1600–1850 cm−1.
The...
The...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...