Related Experiment Video
Updated: Aug 10, 2025

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Fashion-Oriented Image Captioning with External Knowledge Retrieval and Fully Attentive Gates.
Nicholas Moratelli1, Manuele Barraco1, Davide Morelli1
1Department of Engineering "Enzo Ferrari", University of Modena and Reggio Emilia, 41125 Modena, Italy.
This study introduces a new transformer model for generating detailed fashion item descriptions. The model uses external textual memory and a novel gate to significantly improve fashion image captioning accuracy.
Area of Science:
- Computer Vision
- Multimedia Processing
- Natural Language Processing
Background:
- Fashion and e-commerce research is growing in computer vision and multimedia.
- Generating fine-grained, accurate natural language descriptions for fashion items is an under-explored challenge.
Purpose of the Study:
- To develop an advanced model for fine-grained fashion image captioning.
- To overcome limitations of existing approaches in generating accurate fashion descriptions.
Main Methods:
- A transformer-based captioning model integrated with external textual memory.
- Utilized k-nearest neighbor (kNN) searches for memory retrieval.
- Implemented cross-attention operations and a novel fully attentive gate for information flow control.
Main Results:
- The proposed model demonstrated superior performance on the fashion captioning dataset (FACAD).
- Experimental validation confirmed the effectiveness of the architectural strategies.
- The method consistently outperformed baseline and state-of-the-art approaches.
Conclusions:
- The developed transformer model with external memory is highly effective for fashion image captioning.
- The novel architectural components significantly enhance description accuracy.
- This research advances the state-of-the-art in fine-grained fashion item description generation.
More Related Videos
08:25Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
07:12Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
Related Concept Videos
Nonconscious Mimicry
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
Photoreceptors and Visual Pathways
Empathy
Accessory Structures of the Eye
Information Processing Approach