Related Experiment Video
Updated: Nov 5, 2025

13:51
Cross-Modal Multivariate Pattern Analysis
Published on: November 9, 2011
20.2K
Improvement of deep cross-modal retrieval by generating real-valued representation
1U & P U. Patel Department of Computer Engineering, Chandubhai S. Patel Institute of Technology, Charotar University of Science and Technology (CHARUSAT), Changa, India.
Peerj. Computer Science
|May 14, 2021
Summary
This study introduces the Improvement of Deep Cross-Modal Retrieval (IDCMR) framework to bridge the data heterogeneity gap in cross-modal retrieval. IDCMR enhances retrieval accuracy by preserving both intra-modal and inter-modal similarity, outperforming existing methods.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Cross-modal retrieval (CMR) facilitates flexible data searching across different modalities.
- A key challenge in CMR is the heterogeneity gap, stemming from diverse statistical properties of multi-modal data.
- Representation learning, creating a common subspace, is a primary approach to address this gap.
Purpose of the Study:
- To propose a novel framework, Improvement of Deep Cross-Modal Retrieval (IDCMR), for enhancing cross-modal retrieval.
- To generate real-valued representations that effectively bridge the heterogeneity gap.
- To preserve both intra-modal and inter-modal similarity within the learned representations.
Main Methods:
- The IDCMR framework employs tailored training models for text and image modalities to maintain intra-modal similarity.
- Inter-modal similarity is preserved by minimizing a modality-invariance loss function.
- Performance is evaluated using the mean average precision (mAP) metric.
Main Results:
- IDCMR demonstrates superior performance compared to state-of-the-art methods in cross-modal retrieval tasks.
- The framework achieved relative improvements of 4% and 2% in mAP for text-to-image and image-to-text retrieval, respectively.
- Experiments were conducted on the MSCOCO and Xmedia datasets.
Conclusions:
- The proposed IDCMR framework effectively addresses the heterogeneity gap in cross-modal retrieval.
- IDCMR significantly improves retrieval accuracy by preserving essential similarities across modalities.
- The method offers a promising advancement for multi-modal data retrieval systems.
Related Concept Videos
Cross Product
431
The cross product is a fundamental concept in vector algebra that is a vector operation on two different vectors to obtain a third vector. Unlike the scalar product, the cross product results in a vector quantity perpendicular to both the original vectors.
The magnitude of the cross product is obtained by multiplying the magnitude of both the vectors and the sine of the angle between them. This means that a larger angle between the vectors will lead to a greater magnitude of the cross product.
The magnitude of the cross product is obtained by multiplying the magnitude of both the vectors and the sine of the angle between them. This means that a larger angle between the vectors will lead to a greater magnitude of the cross product.
431
Long-term Potentiation
3.0K
Long-term potentiation, or LTP, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTP is the process of synaptic strengthening that occurs over time between pre and postsynaptic neuronal connections. The synaptic strengthening of LTP works in opposition to the synaptic weakening of long-term depression (LTD) and together are the main mechanisms that underlie learning and memory.
Hebbian LTP
LTP can occur when...
Hebbian LTP
LTP can occur when...
3.0K
Long-term Potentiation
56.8K
Long-term potentiation, or LTP, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTP is the process of synaptic strengthening that occurs over time between pre- and postsynaptic neuronal connections. The synaptic strengthening of LTP works in opposition to the synaptic weakening of long-term depression (LTD) and together are the main mechanisms that underlie learning and memory.
56.8K
Sensory Modalities
2.7K
Sensation typically is the process by which the sensory receptors and sense organs detect stimuli from the internal and external environment and transmit this information to the central nervous system for processing.
General senses refer to the broad category of sensory information detected by receptors in the body and can be further grouped into somatic and visceral senses. Somatic sensations include touch, pressure, temperature, and pain and are essential for navigating our environment and...
General senses refer to the broad category of sensory information detected by receptors in the body and can be further grouped into somatic and visceral senses. Somatic sensations include touch, pressure, temperature, and pain and are essential for navigating our environment and...
2.7K
Upsampling
378
Managing signal sampling rates is essential in digital signal processing to maintain signal integrity. A decimated signal, characterized by a reduced frequency range due to its lower sampling rate, can be upsampled by inserting zeros between each sample. This upsampling process expands the original spectrum and introduces repeated spectral replicas at intervals dictated by the new Nyquist frequency. To refine this zero-inserted sequence, it is passed through a lowpass filter with a cutoff...
378
Deconvolution
370
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
370

