Related Experiment Video
Updated: Sep 13, 2026

Coordinate Mapping of Hyolaryngeal Mechanics in Swallowing
Published on: May 6, 2014
Evaluating Intrarater Reliability of Nasopharyngoscopy Ratings Made by Speech-Language Pathologists in a Clinical
Jessica L Chee-Williams1,2, Thomas J Sitzman1,3,4, Adriane Baylis5,6
1Phoenix Children's Center for Cleft and Craniofacial, Division of Plastic Surgery, Phoenix Children's Hospital, AZ.
Purpose:
This study aimed to evaluate intrarater reliability for nasopharyngoscopy ratings of velopharyngeal closure across two contexts: ratings made in a clinical setting and ratings made in a research setting using standardized rating definitions.
Method:
Five speech-language pathologists participated in this study. Two assessments of intrarater reliability were performed: between ratings from clinical reports and re-ratings made in a research setting and videos re-rated twice in a research setting. Nasopharyngoscopy ratings analyzed included velopharyngeal closure pattern, total percent closure, and extent of velar movement. Reliability was calculated using Cohen's kappa and the intraclass correlation coefficient (ICC).
Results:
Intrarater reliability between clinical reports and re-ratings made in a research setting varied for closure pattern from weak (k = .24) to perfect (k = 1.00), for total percent closure from poor (ICC = -.09) to good (ICC = .81), and for extent of velar movement from moderate (ICC = .57) to good (ICC = .80). Compared to intrarater reliability between clinical and research settings, intrarater reliability for videos re-rated twice in a research setting was higher for all raters. For closure pattern, intrarater reliability within a research setting ranged from substantial (k = .68) to perfect (k = 1.00), for total percent closure from moderate (ICC = .70) to excellent (ICC = .99), and for extent of velar movement from good (ICC = .80) to excellent (ICC = .91).
Conclusions:
Findings suggest that nasopharyngoscopy ratings made in a clinical setting may be systematically different than ratings made in a research setting; thus, research results that depend on nasopharyngoscopy ratings may not be generalizable to clinical care. This may be caused by a variety of patient-specific and practice-specific factors, in addition to a lack of standardization. Findings suggest that high reliability for nasopharyngoscopy ratings can be achieved if raters receive consistent training and are following the same rating scale definitions.