Related Experiment Video
Updated: Aug 6, 2026

Pioneering Patient-Specific Approaches for Precision Surgery Using Imaging and Virtual Reality
Published on: April 5, 2024
Scene graph-guided uncertainty decomposition improves confidence calibration in surgical visual question answering
Junzhuo Song1, Yaoxue Xu2, Zihan Zhu2
1College of Water Resources and Hydropower, Sichuan Agricultural University, Yaan, China.
Abstract:
Reliable confidence estimation is essential for surgical visual question answering because overconfident errors may compromise clinical decision support in high-risk settings. However, existing methods typically model predictive uncertainty as a single global quantity, which fails to capture the hierarchical uncertainty arising from ambiguous objects, uncertain tool-tissue interactions, and complex scene context. To address this limitation, we propose a scene graph-guided framework that decomposes uncertainty into object-level, relation-level, and scene-level components and adaptively fuses them for confidence-aware surgical visual question answering. The framework further incorporates Dirichlet-based calibration to improve probabilistic quality and support selective answering. Experiments on the Surgical Scene Graph-Question Answering (SSG-QA) benchmark show that the proposed method achieves an overall accuracy of 63.58%, an Expected Calibration Error of 16.97%, and a Risk-Coverage AUC of 0.0861. Our SG-UD framework achieves competitive calibration performance, significantly improving Brier Score and NLL over baselines, while providing more granular uncertainty interpretability. Performance improves across all question types, with the largest gain observed for relation-type questions (+0.92% over MCAN). Ablation studies further show that uncertainty decomposition contributes most to answer accuracy (+0.70%), whereas Dirichlet calibration is critical for improving probabilistic quality. These findings indicate that explicitly modeling uncertainty across semantic levels can enhance the reliability and interpretability of surgical visual question answering and provide fine-grained confidence information for safer artificial intelligence-assisted surgical decision support.