Related Experiment Video
Updated: Oct 18, 2025

09:41
Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping
Published on: April 21, 2023
1.8K
CrossFuNet: RGB and Depth Cross-Fusion Network for Hand Pose Estimation
Xiaojing Sun1, Bin Wang1, Longxiang Huang2
1College of Information, Mechanical and Electrical Engineering, Shanghai Normal University, Shanghai 200234, China.
Sensors (Basel, Switzerland)
|September 28, 2021
Summary
This study introduces CrossFuNet, a novel network combining RGB and depth data for accurate 3D hand pose estimation. The fusion approach overcomes limitations of individual modalities, improving overall performance.
Area of Science:
- Computer Vision
- Machine Learning
- Robotics
Background:
- RGB-based 3D hand pose estimation faces challenges like self-occlusion and depth ambiguity.
- Depth-based methods are limited by distance sensitivity and indoor-only applicability.
- Combining RGB and depth data offers a promising approach to mitigate individual modality weaknesses.
Purpose of the Study:
- To develop an effective fusion network for enhanced 3D hand pose estimation.
- To address the limitations of existing RGB and depth-based hand pose estimation techniques.
- To introduce a novel fusion module for integrating multi-modal data.
Main Methods:
- A novel RGB and depth information fusion network, CrossFuNet, is proposed.
- Input RGB images and depth maps are processed by separate subnetworks.
- A unique fusion module combines features from both modalities before regressing 3D key-points via heatmaps.
Main Results:
- The CrossFuNet model was validated on two public datasets.
- Experimental results demonstrate superior performance compared to state-of-the-art methods.
- The proposed fusion approach effectively improves the accuracy of 3D hand pose estimation.
Conclusions:
- The CrossFuNet model successfully integrates RGB and depth information for robust 3D hand pose estimation.
- The novel fusion strategy significantly enhances accuracy, overcoming limitations of single-modality approaches.
- This work advances the field of 3D hand pose estimation with a more practical and accurate solution.

