Related Experiment Video
Updated: Jan 16, 2026

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Mutual contextual relation-guided dynamic graph networks for cross-modal image-text retrieval.
G Sucharitha1, B J D Kalyani2, Akella S Narasimha Raju3
1Department of Computer Science and Engineering, Anurag University, Hyderabad, Telangana, India. sucharithasu@gmail.com.
This study introduces a novel dynamic graph network for cross-modal retrieval, enhancing image-text matching by modeling mutual contextual relations. The approach significantly improves precision and recall in retrieving semantically relevant content across modalities.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Cross-modal retrieval is crucial for multimedia search and recommendation due to the rise of multimodal data.
- Challenges include the heterogeneity and semantic gap between image and text representations.
- Existing models often struggle with static feature alignment and inadequate modeling of contextual relationships.
Purpose of the Study:
- To propose a novel mutual contextual relation-guided dynamic graph network for unified and interpretable multimodal representation.
- To enhance image-text matching by dynamically aligning visual and textual features.
- To overcome limitations of existing cross-modal retrieval methods.
Main Methods:
- Integration of Vision Transformer (ViT), BERT, and Graph Convolutional Neural Networks (GCNN).
- Construction of a dynamic cross-modal feature graph (DCMFG) with nodes representing image and text features.
- Dynamic edge updates based on mutual contextual relations (KNN) and an attention-guided mechanism for adaptive alignment.
Main Results:
- Significant performance improvements in precision and recall on benchmark datasets (MirFlickr-25K, NUS-WIDE).
- Demonstrated effectiveness over state-of-the-art methods in cross-modal retrieval.
- Improved interpretability by revealing interactions between image regions and text features.
Conclusions:
- The proposed dynamic graph network effectively addresses the challenges in cross-modal retrieval.
- The method provides a robust and interpretable approach for multimodal representation learning.
- Validated effectiveness for accurate and semantically relevant cross-modal retrieval.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
ER Retrieval Pathway
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...