Related Experiment Video
Updated: Aug 16, 2026

Behavioral Assessment of Hearing in 2 to 4 Year-old Children: A Two-interval, Observer-based Procedure Using Conditioned Play-based Responses
Published on: January 23, 2017
Pediatric Auditory-Perceptual Evaluation of Voice Part 2: A Protocol
Robert Brinton Fujiki1, Sijia Huang2, Rebecca Johnson3
1Department of Otolaryngology-Head and Neck Surgery, Indiana University School of Medicine, Indianapolis.
Purpose:
Although auditory-perceptual voice assessments are crucial for evaluating pediatric voices, no protocol designed for use with children has been established. This study created and validated the Pediatric Auditory-Perceptual Evaluation of Voice (PAPE-V) protocol for use across childhood. Influence of scale type and speaker demographics on voice rating reliability was considered.
Method:
Ten speech-language pathologists (SLPs) specialized in pediatric voice rated voice samples from 102 children aged 3-17 years with and without diagnosed voice disorders (47% female). SLPs were asked to rate each voice for overall voice quality, breathiness, roughness, strain, pitch, and loudness. Voices were rated twice across two separate sessions, with 20% of voices repeated to allow for the calculation of intrarater reliability. In each session, either visual analog scales (VASs) or Likert scales were utilized (order counterbalanced across raters). Interrater and intrarater reliabilities were compared across rating scales and child demographics using intraclass correlation coefficients (ICCs) and Krippendorff's alphas. SLPs' feedback regarding both scale types was also obtained.
Results:
Interrater and intrarater reliabilities were high across all conditions. For both scale types, ICCs were highest for overall voice quality (Likert ICC = .97, VAS ICC = .96) and lowest for pitch (Likert ICC = .82, VAS ICC = .78). Interrater reliability was significantly greater for children presenting with voice disorders than for children with healthy voices (p < .01). Interrater reliability also varied with age, and ratings for children younger than 6 years had significantly better interrater reliability than ratings for children 7-17 years old (p < .001). No significant differences in reliability were observed across scale type or sex. Clinicians indicated that Likert scales were more time efficient than VASs; however, either scale type could be used in the PAPE-V.
Conclusions:
The PAPE-V can be used to perform reliable auditory-perceptual voice assessments in children using either VAS or Likert scales. This protocol constitutes an important step in promoting the reliability and efficiency in pediatric voice assessments.
