Related Experiment Video
Updated: May 28, 2025

Cross-Modal Multivariate Pattern Analysis
Published on: November 9, 2011
Cross modal recipe retrieval with fine grained modal interaction
Fan Zhao1, Yuqing Lu2, Zhuo Yao2
1Faculty of Printing, Packaging Engineering and Digital Media Technology &State Key Laboratory of Eco-hydraulics in Northwest Arid Region, Xi'an University of Technology, Xi'an, 710048, China. vcu@xaut.edu.cn.
Abstract:
Designing systems capable of finding relevant cooking recipes given a user-submitted food image, or vice versa, cross-modal recipe retrieval has gained significant attention in recent years. Numerous advanced techniques have been employed to improve the performance of cross-modal recipe retrieval on general benchmarks. However, leveraging the fine-grained modalities interaction for enhancing multi-modal representation is still limited. Preceding a hierarchical recipe Transformer for encoding individual recipe components, we introduce the cross-component multiscale recipe enriching (CCMRE) module, which enhances the components of the recipe through fully convolutional operations with convolutional kernels of different lengths. Further, we embed a text-contextualized visual enhancing (TCVE) module into an intermediate layer of the image encoder to enrich the visual encoder. Utilizing the similarity between image local features and intermediate recipe representations, TCVE enhances visual representation by deeper model relearning. We conduct a thorough analysis and ablation studies to validate the proposed method, FMI (Fine-grained Modalities Interaction for Cross-Modal Recipe Retrieval). As a result, our method outperforms current SoTA across all metrics on the Recipe1M dataset. Specifically, compared to baseline model, the improvements of + 17.4 R@1 and + 20.5R@1 on the 1 k and 10 k test sets are achieved respectively.
More Related Videos
06:43The Crossmodal Congruency Task as a Means to Obtain an Objective Behavioral Measure in the Rubber Hand Illusion Paradigm
Published on: July 26, 2013
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Related Concept Videos
Sensory Modalities
General senses refer to the broad category of sensory information detected by receptors in the body and can be further grouped into somatic and visceral senses. Somatic sensations include touch, pressure, temperature, and pain and are essential for navigating our environment and...
Optimal Foraging
Crossover Experiments
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
Cross-reactivity
Taste Buds and Receptors
Woodward–Hoffmann Selection Rules and Microscopic Reversibility