Related Experiment Video
Updated: May 5, 2026

05:12
Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
Published on: August 12, 2021
1.9K
A State Space Model for Multiobject Full 3-D Information Estimation From RGB-D Images.
IEEE Transactions on Cybernetics
|March 19, 2025
Summary
This study introduces a novel single-shot method using state space models (SSMs) for efficient and accurate 3-D object understanding from RGB-D images, improving robotic applications.
Area of Science:
- Computer Vision
- Robotics
- Machine Learning
Background:
- Accurate 3-D object understanding is crucial for robotics, autonomous navigation, and augmented reality.
- Current methods often lack end-to-end efficiency and accuracy in predicting full 3-D object information.
Purpose of the Study:
- To develop a single-shot, end-to-end method for predicting the complete 3-D information (pose, size, shape) of multiple objects from a single RGB-D image.
- To leverage State Space Models (SSMs) for enhanced long-range semantic encoding and 3-D inference.
Main Methods:
- A unified model encodes RGB and depth images, integrating them into a latent representation processed by a modified SSM.
- The SSM infers 3-D information via separate heatmap/detection and 3-D information heads.
- An SSM-based shape autoencoder learns canonical shape codes from 3-D point cloud data.
Main Results:
- The proposed method achieves state-of-the-art performance on REAL275, CAMERA25, and Wild6D datasets.
- Significant improvements were observed on the Wild6D dataset, outperforming competitors by 2.6% (IOU-50) and 5.1% (5°10 cm).
Conclusions:
- The end-to-end framework, modified SSM block, and SSM-based shape autoencoder represent significant contributions.
- The method demonstrates superior efficiency and accuracy in 3-D object pose, size, and shape prediction.

