Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Nonconscious Mimicry01:13

Nonconscious Mimicry

4.6K
Nonconscious mimicry occurs when individuals alter their mannerisms to match the behaviors and expressions of those nearby, without intention.
4.6K
Encoding01:19

Encoding

224
Information enters the brain through encoding, which is the input of information into the memory system. Once sensory information is received from the environment, the brain labels or codes it. The information is then organized with similar information and connected to existing concepts. Encoding occurs through automatic processing and effortful processing.
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
224
Photoreceptors and Visual Pathways01:22

Photoreceptors and Visual Pathways

6.2K
At the molecular level, visual signals trigger transformations in photopigment molecules, resulting in changes in the photoreceptor cell's membrane potential. The photon's energy level is denoted by its wavelength, with each specific wavelength of visible light associated with a distinct color. The spectral range of visible light, classified as electromagnetic radiation, spans from 380 to 720 nm. Electromagnetic radiation wavelengths exceeding 720 nm fall under the infrared category,...
6.2K
Empathy02:34

Empathy

9.6K
Some researchers suggest that altruism operates on empathy. Empathy is the capacity to understand another person’s perspective, to feel what he or she feels. An empathetic person makes an emotional connection with others and feels compelled to help (Batson, 1991). Empathy can be expressed in several ways, including cognitive, affective, and motor. 
9.6K
Accessory Structures of the Eye01:17

Accessory Structures of the Eye

1.7K
Optical perception, or vision, is an extraordinary sense dependent on converting light signals received via the ocular organs. These organs, known as eyes, are securely positioned within the bony cavities of the skull, called orbits. The orbits serve a dual purpose: a protective shield for the ocular globes and a stable attachment point for the soft ocular tissues. The eye's external protective mechanisms include the eyelids, which are edged with lashes that act as a barrier against foreign...
1.7K
Information Processing Approach01:30

Information Processing Approach

90
The information-processing theory of cognitive development centers on fundamental mental processes, including attention, memory, and problem-solving skills. Researchers in this field examine how cognitive abilities, such as working memory, evolve and influence children's overall development. Studies indicate that children with stronger working memory tend to excel in reading comprehension, math, and problem-solving compared to peers with less efficient memory skills. Low working memory is...
90

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Computer Vision in Human Analysis: From Face and Body to Clothes.

Sensors (Basel, Switzerland)·2023
Same author

A Systematic Comparison of Depth Map Representations for Face Recognition.

Sensors (Basel, Switzerland)·2021
Same author

Driver Face Verification with Depth Maps.

Sensors (Basel, Switzerland)·2019
See all related articles

Related Experiment Video

Updated: Aug 10, 2025

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
07:36

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects

Published on: November 30, 2018

15.8K

Fashion-Oriented Image Captioning with External Knowledge Retrieval and Fully Attentive Gates.

Nicholas Moratelli1, Manuele Barraco1, Davide Morelli1

  • 1Department of Engineering "Enzo Ferrari", University of Modena and Reggio Emilia, 41125 Modena, Italy.

Sensors (Basel, Switzerland)
|February 11, 2023
PubMed
Summary

This study introduces a new transformer model for generating detailed fashion item descriptions. The model uses external textual memory and a novel gate to significantly improve fashion image captioning accuracy.

Keywords:
fashion captioningimage captioningknowledge retrievalvision-and-language

More Related Videos

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

9.1K
Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
07:12

Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss

Published on: April 11, 2025

489

Related Experiment Videos

Last Updated: Aug 10, 2025

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
07:36

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects

Published on: November 30, 2018

15.8K
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

9.1K
Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
07:12

Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss

Published on: April 11, 2025

489

Area of Science:

  • Computer Vision
  • Multimedia Processing
  • Natural Language Processing

Background:

  • Fashion and e-commerce research is growing in computer vision and multimedia.
  • Generating fine-grained, accurate natural language descriptions for fashion items is an under-explored challenge.

Purpose of the Study:

  • To develop an advanced model for fine-grained fashion image captioning.
  • To overcome limitations of existing approaches in generating accurate fashion descriptions.

Main Methods:

  • A transformer-based captioning model integrated with external textual memory.
  • Utilized k-nearest neighbor (kNN) searches for memory retrieval.
  • Implemented cross-attention operations and a novel fully attentive gate for information flow control.

Main Results:

  • The proposed model demonstrated superior performance on the fashion captioning dataset (FACAD).
  • Experimental validation confirmed the effectiveness of the architectural strategies.
  • The method consistently outperformed baseline and state-of-the-art approaches.

Conclusions:

  • The developed transformer model with external memory is highly effective for fashion image captioning.
  • The novel architectural components significantly enhance description accuracy.
  • This research advances the state-of-the-art in fine-grained fashion item description generation.