Related Experiment Video
Updated: Aug 27, 2026

Methodology for Establishing a Community-Wide Life Laboratory for Capturing Unobtrusive and Continuous Remote Activity and Health Data
Published on: July 27, 2018
Online, Crowdsourced Sampling (OCS) Platforms for Large-Scale Data Collection in Voice Disorders Research: Prices,
Christopher S Apfelbach1, Lady Catherine Cantor-Cutiva2, Eric J Hunter3
1Department of Speech-Language-Hearing Sciences, The University of Minnesota, Minneapolis, Minnesota.
Objective:
Online, crowdsourced sampling (OCS) platforms such as Amazon's Mechanical Turk and CloudResearch Connect enable faster, cheaper, and larger-scale data collection than in-person sampling methods. However, OCS platforms are often criticized due to data quality concerns. This study examines a battery of screening tools to determine which best predicts data quality and to quantify the influence of low-quality data on the statistical properties of the samples. Finally, we issue recommendations to researchers interested in OCS-based clinical voice research that may reduce costly missteps when fielding their first OCS studies.
Methods:
To evaluate the suitability of OCS for clinical voice research, survey-based measures of vocal function, vocal fatigue, personality, communicative quality of life, and other factors were collected on Qualtrics from (1) a convenience sample of undergraduate students (n = 47), (2) Mechanical Turk users (n = 495), and (3) Connect users (n = 99) over 6 months. Low-quality responses were eliminated using six screening tools in four nested stages. Logistic regression was used to identify variables that strongly predicted data quality. Finally, patient-reported outcome measure (PROM) responses were compared between the high- and low-quality groups.
Results:
Depending on the stage, data quality screenings flagged between 5.00% (n = 32) and 13.6% (n = 87) of responses as low-quality. Hispanic/Latino ethnicity and long response times most strongly predicted low-quality responses, which uniformly exhibited more severe ratings of health- and voice-related disability than did high-quality responses.
Conclusions:
Content-based screenings, particularly analysis of responses to open-ended questions, meaningfully augmented automated screenings. Low-quality respondents systematically over-reported health- and voice-related disability, a severity bias with plausible economic and behavioral explanations. The emergence of AI-generated responses represents an evolving and distinct challenge that content-based screening may be increasingly insufficient to address. We provide recommendations for collecting high-quality data on OCS platforms to minimize barriers to entry for future clinical voice research.
Related Concept Videos
Convenience Sampling Method
Convenience sampling is a non-random method of sample selection; this method selects individuals that are easily accessible and may result in biased data. For example, a marketing...
Surveys
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Systematic Sampling Method
Systematic sampling is one of the simplest methods...
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...