Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Spatially Aware Pair Proposal for Panoptic Scene Graph Generation.

Sensors (Basel, Switzerland)·2026
Same author

A Cartesian-Based Trajectory Optimization with Jerk Constraints for a Robot.

Entropy (Basel, Switzerland)·2023
See all related articles

Related Experiment Video

Updated: Aug 5, 2025

Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping
09:41

Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping

Published on: April 21, 2023

1.7K

Bilateral Cross-Modal Fusion Network for Robot Grasp Detection.

Qiang Zhang1,2, Xueying Sun1,2

  • 1School of Automation, Jiangsu University of Science and Technology, No. 666 Changhui Road, Zhenjiang 212100, China.

Sensors (Basel, Switzerland)
|March 30, 2023
PubMed
Summary

This study introduces a novel tri-stream fusion architecture for visual grasp detection, enhancing robot accuracy by integrating RGB and depth data. The method achieves high detection rates, improving robotic grasping capabilities.

Keywords:
channel interactioncross-modality fusionrobot grasp detection

More Related Videos

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

592
Investigating Object Representations in the Macaque Dorsal Visual Stream Using Single-unit Recordings
07:08

Investigating Object Representations in the Macaque Dorsal Visual Stream Using Single-unit Recordings

Published on: August 1, 2018

8.4K

Related Experiment Videos

Last Updated: Aug 5, 2025

Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping
09:41

Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping

Published on: April 21, 2023

1.7K
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

592
Investigating Object Representations in the Macaque Dorsal Visual Stream Using Single-unit Recordings
07:08

Investigating Object Representations in the Macaque Dorsal Visual Stream Using Single-unit Recordings

Published on: August 1, 2018

8.4K

Area of Science:

  • Robotics
  • Computer Vision
  • Artificial Intelligence

Background:

  • Accurate target pose estimation using RGB and depth data is crucial for vision-based robot grasping.
  • Existing methods face challenges in effectively fusing multimodal information for precise grasp detection.

Purpose of the Study:

  • To propose a novel tri-stream cross-modal fusion architecture for 2-DoF visual grasp detection.
  • To enhance the aggregation of multiscale information from RGB and depth data.
  • To improve the accuracy and robustness of robot grasping systems.

Main Methods:

  • Developed a tri-stream cross-modal fusion architecture integrating RGB and depth information.
  • Introduced a novel modal interaction module (MIM) utilizing spatial-wise cross-attention.
  • Employed channel interaction modules (CIM) for enhanced cross-modal feature aggregation.
  • Utilized a hierarchical structure with skipping connections for multiscale information aggregation.

Main Results:

  • Achieved 99.4% image-wise detection accuracy on the Cornell dataset and 96.7% on the Jacquard dataset.
  • Reached 97.8% object-wise detection accuracy on the Cornell dataset and 94.6% on the Jacquard dataset.
  • Demonstrated a 94.5% success rate in physical experiments using a 6-DoF Elite robot.

Conclusions:

  • The proposed tri-stream cross-modal fusion architecture significantly improves visual grasp detection accuracy.
  • The novel MIM and CIM modules effectively capture and aggregate cross-modal and multiscale features.
  • The method shows superior performance in both dataset evaluations and real-world robotic grasping tasks.