Related Experiment Video
Updated: Aug 28, 2026

Minimally Invasive Murine Laryngoscopy for Close-Up Imaging of Laryngeal Motion During Breathing and Swallowing
Published on: December 1, 2023
Human Evaluation of Synthetic Videostroboscopic Laryngeal Images Generated Using StyleGAN3
Danielle A Morrison1, Abhinita S Mohanty2, Lucian Sulica1
1Department of Otolaryngology-Head and Neck Surgery, Sean Parker Institute for the Voice, Weill Cornell Medicine, New York, New York, USA.
Objective:
To evaluate the perceptual realism of synthetic videostroboscopic laryngeal images and determine the relationship between training duration and clinician classification accuracy.
Methods:
Synthetic images were generated using StyleGAN3 from two age-stratified datasets: Dataset A (≥ 65 years) and Dataset B (< 65 years). A total of 114 clinicians evaluated a randomized, balanced-controlled 36-image survey drawn from a 144-image pool containing real frames and five StyleGAN3 training intervals (5120-25,000 kimg). Accuracy was analyzed across image classes, experience levels, and viewing devices.
Results:
Clinicians demonstrated significantly higher mean accuracy identifying true clinical frames than synthetic generations (70.9% vs. 57.8%, p < 0.001). Overall survey accuracy was 59.9% ± 14.2%, with realism peaking at 20,120 kimg (44.8% accuracy). Accuracy dropped significantly between 5120 kimg (79.6%) and 10,120 kimg (59.1%, p < 0.001), with diminishing returns thereafter. No significant accuracy difference existed between Datasets A and B (p = 0.27). The effect of specialty reached borderline significance (p = 0.059), though limited by severe subgroup imbalances (e.g., n = 5 fellows, n = 2 residents). Computer users (63.9%) were significantly more accurate than phone users (56.2%, p = 0.014). A confounder analysis confirmed specialty and device choice were statistically independent (p = 0.164).
Conclusion:
StyleGAN3 produces synthetic laryngeal images demonstrating high perceptual realism. The significant real-versus-synthetic accuracy gap highlights that while true frames remain distinct, synthetic generations achieve profound ambiguity. Perceptual realism plateaus early (10,120-20,120 kimg), suggesting moderate training durations can produce human-perceived realism while reducing computational and environmental demands.
Level Of Evidence:
N/A.
