Related Experiment Video
Updated: Aug 5, 2025

09:41
Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping
Published on: April 21, 2023
1.7K
Bilateral Cross-Modal Fusion Network for Robot Grasp Detection.
Qiang Zhang1,2, Xueying Sun1,2
1School of Automation, Jiangsu University of Science and Technology, No. 666 Changhui Road, Zhenjiang 212100, China.
Sensors (Basel, Switzerland)
|March 30, 2023
Summary
This study introduces a novel tri-stream fusion architecture for visual grasp detection, enhancing robot accuracy by integrating RGB and depth data. The method achieves high detection rates, improving robotic grasping capabilities.
Area of Science:
- Robotics
- Computer Vision
- Artificial Intelligence
Background:
- Accurate target pose estimation using RGB and depth data is crucial for vision-based robot grasping.
- Existing methods face challenges in effectively fusing multimodal information for precise grasp detection.
Purpose of the Study:
- To propose a novel tri-stream cross-modal fusion architecture for 2-DoF visual grasp detection.
- To enhance the aggregation of multiscale information from RGB and depth data.
- To improve the accuracy and robustness of robot grasping systems.
Main Methods:
- Developed a tri-stream cross-modal fusion architecture integrating RGB and depth information.
- Introduced a novel modal interaction module (MIM) utilizing spatial-wise cross-attention.
- Employed channel interaction modules (CIM) for enhanced cross-modal feature aggregation.
- Utilized a hierarchical structure with skipping connections for multiscale information aggregation.
Main Results:
- Achieved 99.4% image-wise detection accuracy on the Cornell dataset and 96.7% on the Jacquard dataset.
- Reached 97.8% object-wise detection accuracy on the Cornell dataset and 94.6% on the Jacquard dataset.
- Demonstrated a 94.5% success rate in physical experiments using a 6-DoF Elite robot.
Conclusions:
- The proposed tri-stream cross-modal fusion architecture significantly improves visual grasp detection accuracy.
- The novel MIM and CIM modules effectively capture and aggregate cross-modal and multiscale features.
- The method shows superior performance in both dataset evaluations and real-world robotic grasping tasks.

