Related Experiment Video
Updated: May 27, 2026

08:21
Utilizing a 3D Printed Laparoscopic Nissen Fundoplication Model to Shorten a Resident's Learning Curve
Published on: August 15, 2025
Evaluating Injection Laryngoplasty Skills Using a Foundation Model: A Feasibility Study
Alex T Cheng1, Abdulla Elkhadrawy1, Sean A Setzen1
1Department of Otolaryngology-Head and Neck Surgery, Weill Cornell Medicine, New York, New York, USA.
The Laryngoscope
|May 26, 2026
Summary
Few-shot prompting with Google Gemini 2.5 Pro successfully assessed surgical skill in simulated injection laryngoplasty, distinguishing expert from trainee performance. Averaging repeat evaluations can mitigate model variability for this promising assessment tool.
Area of Science:
- Artificial Intelligence in Medicine
- Surgical Skill Assessment
- Medical Simulation
Background:
- Assessing surgical proficiency is critical for patient safety and training effectiveness.
- Objective and reliable methods for evaluating procedural skills are needed.
- Multimodal foundation models offer potential for automated performance analysis.
Purpose of the Study:
- To evaluate the construct validity of Google Gemini 2.5 Pro for assessing simulated injection laryngoplasty.
- To compare zero-shot versus few-shot prompting strategies for skill assessment.
- To determine model reliability and stability in performance evaluation.
Main Methods:
- Thirty simulated injection laryngoplasty videos were stratified by operator experience (novice, intermediate, expert).
- Google Gemini 2.5 Pro evaluated videos using zero-shot and few-shot prompting strategies.
- Model performance was compared against operator experience, with reliability assessed via 90 repeated trials.
Main Results:
- Zero-shot prompting failed to discriminate between skill levels (Spearman's ρ = 0.12, p = 0.52).
- Few-shot prompting showed strong correlation with experience (Spearman's ρ = 0.66, p = 0.0002) and stratified skill levels.
- Few-shot model significantly differentiated experts from novices and intermediates, improving precision and reducing error.
Conclusions:
- General-purpose multimodal models require calibration (e.g., few-shot prompting) for surgical judgment.
- Few-shot prompting effectively calibrated Gemini 2.5 Pro to distinguish expert from trainee performance.
- Model variability necessitates mitigation strategies, such as averaging repeated evaluations, for scalable assessment.

