General 3D Vision-Language Model With Fast Rendering and Pre-Training Vision-Language Alignment

Summary

This study introduces WS3D++, a framework for 3D scene understanding that excels with limited labels. It enables open-vocabulary recognition and achieves state-of-the-art performance in semantic and instance segmentation for 3D point clouds.