Related Experiment Video
Updated: Aug 26, 2026

Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
Published on: August 12, 2021
Why Only 3D for 3D? 2D-Guided Supervision for 3D Medical Vision-Language Models
Qilong Zhao1, Ziyuan Qin1, Liang Zhao1
1Department of Computer Science, Emory University, Atlanta, GA, USA.
Abstract:
Training 3D medical vision-language models (VLMs) requires paired volumetric scans and expert annotations such as reports or segmentation masks, which are costly and limited. Meanwhile, large collections of unlabeled 3D scans and strong 2D medical VLMs are increasingly available. We study a simple paradigm: using pretrained 2D models as automatic annotators to supervise 3D VLMs. A 2D teacher generates slice-level descriptions or masks that are aggregated into volume-level pseudo-labels to train a 3D student operating on full volumes at inference. We instantiate this 2D-to-3D supervision framework for report generation and segmentation. In label-scarce regimes, 2D-derived pseudo-labels significantly improve data efficiency. For report generation, pseudo-reports outperform or match ground truth when reports are limited and remain competitive at larger scales. For segmentation, pseudo-masks enable learning from unlabeled volumes and further improve performance when combined with limited expert masks. Overall, strong 2D models can act as scalable annotators that complement scarce 3D labels. This data-centric perspective offers a practical path to scaling 3D medical VLMs by leveraging abundant 2D expertise without changing model architectures.
