Related Experiment Video
Updated: Jan 9, 2026

Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility
Published on: August 9, 2024
ViTCAI: A Vision-Language Model for Automated Triage and Captioning of Postoperative Incision Images Submitted by
Abstract:
Surgical site infections (SSIs) are common postoperative complications that increase patient morbidity, hospital stays, and healthcare costs. Early detection and precise documentation are critical for timely intervention and improved outcomes. While patients often submit images of their wounds through a patient portal for remote monitoring, manual review is challenging due to image variability, high volume, and subjectivity, underscoring the need for automated assessment tools. In this paper, we present Vision Transformer for CAptioning of Incision Images (ViTCAI), a vision-language model designed for automated triage and captioning of postoperative incision images. Fine-tuned on a clinically annotated dataset, ViTCAI improves descriptive accuracy in identifying surgical incisions and SSIs. Our results show that ViTCAI provides consistent, detailed captions that can support clinical decision-making, reducing workload and enhancing diagnostic efficiency in postoperative care.

