Related Experiment Video
Updated: Jul 6, 2026

12:08
From Voxels to Knowledge: A Practical Guide to the Segmentation of Complex Electron Microscopy 3D-Data
Published on: August 13, 2014
25.0K
End-to-end Fusion3DGS: label-efficient multi-modal 3D instance segmentation based on Gaussian splatting.
Zhiyuan Wang1,2, Chuang Luo1,2, Jianping Zhang3,4
1School of Transportation and Logistics, Southwest Jiaotong University, Chengdu, 610031, China.
Scientific Reports
|January 6, 2026
Summary
Fusion3DGS enables efficient 3D instance segmentation using only 2D masks, reducing the need for costly 3D data. This novel framework enhances perception in embodied systems by leveraging readily available RGB images.
Area of Science:
- Computer Vision
- Robotics
- Machine Learning
Background:
- Accurate 3D instance segmentation is crucial for embodied systems but relies on expensive 3D point cloud annotations.
- Current methods face challenges in scalability across different sensors and environments due to data acquisition costs.
Purpose of the Study:
- To develop a label-efficient framework for 3D instance segmentation using readily available 2D data.
- To address the bottleneck of dense 3D annotations in current perception systems.
Main Methods:
- Introduced Fusion3DGS, an end-to-end framework coupling 3D Gaussian Splatting with 2D-3D neural processing.
- Utilized multi-view RGB images with 2D instance masks to optimize a compact anisotropic Gaussian scene representation.
- Employed an occlusion-aware cross-attention fusion stack for instance reasoning, incorporating a weight-sharing lock and rendering consistency objective.
Main Results:
- Achieved label-efficient 3D instance segmentation from 2D instance masks, significantly reducing annotation requirements.
- Demonstrated improved boundary fidelity under occlusion and varying viewpoints through rendering consistency.
- Validated the framework's practicality for large-scale deployment using widely available RGB data.
Conclusions:
- Fusion3DGS offers a practical solution for 3D instance segmentation by overcoming the limitations of dense 3D data.
- The framework's ability to learn from 2D supervision makes it scalable and adaptable for diverse embodied AI applications.

