Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Collisions in Multiple Dimensions: Problem Solving01:06

Collisions in Multiple Dimensions: Problem Solving

5.3K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
5.3K
ER Retrieval Pathway01:45

ER Retrieval Pathway

4.7K
In the secretory pathway, vesicles transport proteins from one cellular compartment to another in forward transport to deliver the protein to its correct location. Occasionally, misfolded proteins and incorrect proteins escape their original compartments, and a retrieval pathway is used to return the escaped proteins to their original compartment.
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
4.7K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Pose Estimation of Unmanned Underwater Vehicles Using Augmented Reality Marker-Based Simulations.

Journal of visualized experiments : JoVE·2026
Same author

HYDRA-XAI dual-backbone disaster scene recognition using ResNet50-Swin transformer feature fusion, explainable evidence, and an operational recommender.

Scientific reports·2026
Same author

ColoXAI-RecomNet: Explainable Recommender Framework for Colorectal Cancer Classification Using Integrated CNN Ensemble and LIME Interpretability.

Journal of imaging informatics in medicine·2026
Same author

Impacts of structural and soil parameters on the seismic response of neighbouring buildings: a numerical investigation.

Scientific reports·2026
Same author

Magnetohydrodynamic peristaltic flow of hybrid nanofluid in an asymmetric channel with thermal radiation and shape factors: an application of coolant systems.

Discover nano·2026
Same author

Hybrid deep feature integration model for robust deepfake detection using transfer-learned neural networks.

Frontiers in artificial intelligence·2026

Related Experiment Video

Updated: Jan 16, 2026

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
07:36

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects

Published on: November 30, 2018

16.3K

Mutual contextual relation-guided dynamic graph networks for cross-modal image-text retrieval.

G Sucharitha1, B J D Kalyani2, Akella S Narasimha Raju3

  • 1Department of Computer Science and Engineering, Anurag University, Hyderabad, Telangana, India. sucharithasu@gmail.com.

Scientific Reports
|October 1, 2025
PubMed
Summary

This study introduces a novel dynamic graph network for cross-modal retrieval, enhancing image-text matching by modeling mutual contextual relations. The approach significantly improves precision and recall in retrieving semantically relevant content across modalities.

Keywords:
BERTCross-modal retrievalDynamic attention mechanismGCNNMultimodal graph networkMutual contextual relationViT

More Related Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.3K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.0K

Related Experiment Videos

Last Updated: Jan 16, 2026

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
07:36

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects

Published on: November 30, 2018

16.3K
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.3K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.0K

Area of Science:

  • Computer Science
  • Artificial Intelligence
  • Machine Learning

Background:

  • Cross-modal retrieval is crucial for multimedia search and recommendation due to the rise of multimodal data.
  • Challenges include the heterogeneity and semantic gap between image and text representations.
  • Existing models often struggle with static feature alignment and inadequate modeling of contextual relationships.

Purpose of the Study:

  • To propose a novel mutual contextual relation-guided dynamic graph network for unified and interpretable multimodal representation.
  • To enhance image-text matching by dynamically aligning visual and textual features.
  • To overcome limitations of existing cross-modal retrieval methods.

Main Methods:

  • Integration of Vision Transformer (ViT), BERT, and Graph Convolutional Neural Networks (GCNN).
  • Construction of a dynamic cross-modal feature graph (DCMFG) with nodes representing image and text features.
  • Dynamic edge updates based on mutual contextual relations (KNN) and an attention-guided mechanism for adaptive alignment.

Main Results:

  • Significant performance improvements in precision and recall on benchmark datasets (MirFlickr-25K, NUS-WIDE).
  • Demonstrated effectiveness over state-of-the-art methods in cross-modal retrieval.
  • Improved interpretability by revealing interactions between image regions and text features.

Conclusions:

  • The proposed dynamic graph network effectively addresses the challenges in cross-modal retrieval.
  • The method provides a robust and interpretable approach for multimodal representation learning.
  • Validated effectiveness for accurate and semantically relevant cross-modal retrieval.