Related Experiment Video
Updated: Jan 31, 2026

Making Sense of Listening: The IMAP Test Battery
Published on: October 11, 2010
Does the Use of Crowdsourced Listeners Yield Different Speech Intelligibility Results Than In-Person Listeners for
Heather D Salvo1, Tristan J Mahr1, Carly Sandgren1
1Waisman Center, University of Wisconsin-Madison.
Purpose:
We examined the performance of crowdsourced listeners compared with in-person listeners on the measurement of speech intelligibility for typically developing children. We used three different in-task quality check criteria to screen listeners and examined between-listener intelligibility differences and interrater reliability under each criterion. We also examined how crowdsourced intelligibility results compared with in-person results.
Method:
Sixty neurotypical children between ages 2;6 and 9;11 (years;months), drawn from Hustad et al. (2021), contributed speech samples. We used the online platform, Prolific, to collect intelligibility data from five crowdsourced listeners per child (N = 300 total) and compared scores with in-person results from two listeners per child. We used intraclass correlation coefficients (ICCs) and computed pairwise differences among listener groups for each of three in-task quality criteria groups and the in-person group. We modeled intelligibility as a function of listener source (in-person vs. crowdsourced) and child age using mixed-effects regression with smoothing splines.
Results:
Lower ICCs and larger between-listener differences were observed for crowdsourced compared to in-person listeners, regardless of in-task quality check criteria, but in-task quality check criteria reduced the disparity. Crowdsourced listeners produced intelligibility scores that were up to 7 percentage points lower than in-person listeners, even under the most stringent in-task quality check criterion. Results showed the same patterns of change with children's age as in-person listener findings. Children with midrange (65%-83%) intelligibilities were the most negatively impacted by the use of crowdsourced listeners.
Conclusions:
Rigorous in-task quality check criteria improved crowdsourced listener data. Speakers with midrange intelligibility were the most negatively impacted by the use of crowdsourced listeners, with an intelligibility difference of about 7 percentage points. Intelligibility data obtained with crowdsourced listeners should be interpreted with caution, and future studies should evaluate how crowdsourced intelligibility data differs from in-person data for disordered populations of speakers.
Supplemental Material:
https://doi.org/10.23641/asha.31101451.
Related Concept Videos
Techniques of therapeutic communication I: Active Listening, Sharing Observations, Validation, and Using Touch
Therapeutic communication is not the same as social interaction. Social interaction has no goal or purpose and consists of casual information sharing, whereas therapeutic communication has a plan or purpose for the conversation. Therapeutic...
ATP Yield
The ETC is embedded in the inner mitochondrial membrane and is comprised of four main protein complexes and an ATP synthase. NADH and FADH2 pass electrons to these complexes, which pump protons into the intermembrane space. This distribution of...
Reaction Yield
Typical Model Studies
Intelligence
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...

