Related Experiment Video
Updated: Aug 29, 2026

Robotic Cochlear Implantation for Direct Cochlear Access
Published on: June 16, 2022
Embodied AI in the operating room: a voice-interactive robotic scrub nurse for surgical instrument handoffs
Mustafa Khan1, Joao Bombardelli2
1UTHealth Houston, Division of Plastic Surgery, 6410 Fannin St, Suite 1400, Houston, TX, 77030, USA.
Abstract:
Embodied artificial intelligence may extend surgical robotics beyond teleoperation by enabling robots to interpret natural-language commands, perceive dynamic environments, and generate adaptive physical actions. Surgical instrument exchange provides a clinically recognizable test case requiring communication, visual localization, dexterity, and bidirectional interaction. We developed a voice-interactive robotic scrub nurse through task-specific adaptation of the pretrained π0.5 vision-language-action model capable of voice-conditioned surgical instrument handoff, take-back, and spoken acknowledgment. A bimanual robotic platform was trained by fine-tuning the pretrained π0.5 model on 350 human-teleoperated demonstrations. The system integrated real-time multicamera observations, Whisper large-v2 speech recognition, language-conditioned robotic action generation, and text-to-speech acknowledgment. Performance was evaluated in 100 autonomous trials across five tasks: needle-driver handoff, forceps handoff, needle-and-suture handoff, needle-driver take-back, and forceps take-back. Each task included 10 trials using spatial configurations represented during training and 10 trials using held-out spatial configurations within the same environment. Outcomes included command interpretation, stage-wise completion, end-to-end success, completion time, and failure mode. Recovery after interference was examined qualitatively in separate representative demonstrations. The system completed 81% of trials successfully. Success was 95% for both take-back tasks, 85% for forceps handoff, 80% for needle-driver handoff, and 50% for needle-and-suture handoff. Performance decreased from 92% in seen layouts to 70% in held-out layouts. Speech-command interpretation succeeded in 99% of trials. Once a secure grasp was achieved, downstream manipulation succeeded in 81 of 82 trials. Missed grasp was the predominant failure mode. A unified learned policy enabled voice-conditioned bidirectional surgical instrument exchange and demonstrated meaningful generalization to held-out layouts within the same environment. Improved grasp acquisition, expanded instrument diversity, quantitative recovery testing, and evaluation in realistic operating-room workflows are required before clinical translation or integration with complete back-table workflows.
