Related Experiment Video
Updated: Aug 9, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Dual-stage framework with soft-label distillation and spatial prompting for image-text retrieval
Ran Jin1,2, Zhengang Li1, Fang Deng1
1Zhejiang Wanli University, School of Big Data and Software Engineering, Ningbo, China.
This study introduces a dual-stage training framework to improve image-text retrieval accuracy by addressing inter-modal matching and fine-grained localization deficiencies using Soft Label Distillation (SLD) and Spatial Text Prompt (STP). The novel approach enhances cross-modal understanding for better retrieval performance.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Vision-language pre-training (VLP) has advanced cross-modal tasks.
- Current image-text retrieval methods suffer from inter-modal matching and fine-grained localization deficiencies, limiting accuracy.
Purpose of the Study:
- To propose a novel dual-stage training framework to overcome limitations in image-text retrieval.
- To enhance the accuracy and fine-grained alignment in cross-modal retrieval tasks.
Main Methods:
- Implemented a dual-stage training framework.
- Utilized Soft Label Distillation (SLD) in the first stage to align image-text contrastive relationships and mitigate overfitting.
- Introduced Spatial Text Prompt (STP) in the second stage to improve visual grounding and fine-grained alignment.
Main Results:
- The proposed method significantly outperforms state-of-the-art approaches on standard image-text retrieval datasets.
- Demonstrated improved accuracy in cross-modal retrieval tasks through enhanced alignment.
Conclusions:
- The dual-stage training framework effectively addresses key challenges in image-text retrieval.
- The combination of SLD and STP leads to superior performance in cross-modal understanding and retrieval.
Related Concept Videos
Leaky Scanning
Fixation and Sectioning
The simplest type of preparation is the wet mount, in which the specimen is placed in a drop of liquid on the slide. A liquid specimen can be directly deposited on the slide using a dropper. Solid specimens, such as skin scraping, can be placed on the slide before adding a drop of liquid to prepare the wet mount. Sometimes the liquid is simply water, but stains are often added...
Methods of Documentation IV: Focus Charting
It typically involves three columns for recording information:
Guidelines and Strategies for Safe Computer Charting
Maintain Confidentiality and Security:
Spanning Openings in Brick Walls
Lintels are primary supports used to span openings and can be crafted from materials such as reinforced concrete, steel-reinforced brick masonry, or simple steel angles. These are straightforward to install and are typically concealed...
Masonry Curtain Walls

