Related Experiment Video
Updated: Jul 10, 2026

Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications
Published on: January 5, 2024
Feasibility of a Pediatric Voice Protocol for Artificial Intelligence Research
Siyu Miao1, Jinny Choi1, Laurie Russell2
1Department of Otolaryngology - Head and Neck Surgery, Hospital for Sick Children, University of Toronto, Toronto, ON, Canada.
Introduction:
There is increasing interest in leveraging artificial intelligence (AI) research in voice for pediatric healthcare. However, creating high-quality databases for pediatric voice AI research presents unique challenges due to developmental differences in attention and adherence. Effective data collection methods must accommodate age-specific needs while ensuring consistency in recording conditions and feasibility across diverse age groups. Research on standardized approaches for capturing pediatric voice data in clinical settings remains limited.
Methods:
This prospective observational study incorporated demographic surveys and validated questionnaires to collect medical and demographic data, alongside voice acoustic tasks grouped into four age-based categories. Data collection was conducted using two iPads with headphones. The primary outcome was protocol feasibility, encompassing proportions of successful task completion, headphones usage, required prompting, and average length of time to complete the full assessment. Secondary outcomes were the fundamental frequency (Hz), jitter (%), shimmer (%), and harmonic-to-noise ratio (dB) of the recordings to measure acoustic quality.
Results:
Hundred participants were recruited (50 females, 50 males) from a tertiary care otolaryngology clinic at the Hospital for Sick Children between October 2024 and January 2025. The majority (75%) completed the assessment within 10-25 minutes, while a smaller proportion required additional time due to young age or medical complexity. Headphone compliance was high (90%), with alternative methods to ensure data continuity. Younger participants required more prompting, and male participants in younger age groups were less likely to tolerate headphones. Acoustic measures were consistent with published age- and sex- based norms.
Conclusion:
This study demonstrates the feasibility of collecting acoustic data to support AI research in pediatric populations. Addressing age-specific needs and parental concerns is critical for optimizing engagement and ensuring high-quality datasets. These findings provide valuable insights for improving recruitment, adherence, and data integrity, contributing to the creation of robust pediatric voice AI databases for future clinical and research applications.

