Related Experiment Video
Updated: Mar 31, 2026

Adapting Human Videofluoroscopic Swallow Study Methods to Detect and Characterize Dysphagia in Murine Disease Models
Published on: March 1, 2015
Assessing the Performance and Reliability of Deep Learning Autosegmentation in Videofluoroscopic Swallowing Studies:
Wei-Kai Chuang1, Bing-Fong Lin2, Yu-Hao Lee3
1Department of Radiation Oncology, Shuang Ho Hospital, Taipei Medical University, New Taipei City, Taiwan; Department of Biomedical Imaging and Radiological Sciences, National Yang Ming Chiao Tung University, Taipei, Taiwan; Department of Radiation Oncology, Saint Paul's Hospital, Taoyuan, Taiwan.
Objective:
To systematically evaluate the accuracy and reliability of deep learning-based autosegmentation methods in videofluoroscopic swallowing study (VFSS) through meta-analysis.
Data Sources:
A comprehensive literature search was conducted across PubMed, IEEE Xplore, Embase, Web of Science, and Cochrane Library databases for studies published in English between 2013 and 2025.
Study Selection:
Studies were included if they applied deep learning techniques to the autosegmentation of anatomical structures in VFSS, specifically the bolus, cervical spine, hyoid bone, or thyroid cartilage-vocal fold complex, and reported quantitative performance metrics such as the Dice similarity coefficient.
Data Extraction:
Two independent reviewers extracted data on study characteristics, segmentation targets, deep learning model types, and performance metrics. Methodological quality was assessed using the Checklist for Artificial Intelligence in Medical Imaging and Quality Assessment of Diagnostic Accuracy Studies-2 tools.
Data Synthesis:
Ten studies met the inclusion criteria. A random-effects meta-analysis yielded an overall pooled Dice score of 0.83 (95% CI, 0.76-0.88; I²=77%). Subgroup analyses showed similar performance for bolus segmentation (pooled Dice score=0.84; 95% CI, 0.70-0.92; I²=74%) and cervical spine segmentation (pooled Dice score=0.83; 95% CI, 0.69-0.91; I²=87%). Despite high accuracy, substantial heterogeneity was observed.
Conclusions:
Deep learning-based autosegmentation in VFSS demonstrates promising accuracy across different anatomical targets. However, methodological variability among studies underscores the need for standardized protocols, multicenter datasets, and comparative evaluations of model architectures to enhance generalizability and clinical utility.

