Related Experiment Video
Updated: Sep 12, 2025

A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
Published on: September 4, 2019
Data quality in crowdsourcing and spamming behavior detection
Yang Ba1, Michelle V Mancenido2, Erin K Chiou3
1Ira A. Fulton Schools of Engineering, School of Computing and Augmented Intelligence, Data Science, Analytics and Engineering, Arizona State University, Suite 342AE, 3rd floor 699 S. Mill Avenue, 85281, Tempe, AZ, USA. yangba@asu.edu.
Abstract:
As crowdsourcing emerges as an efficient and cost-effective method for obtaining labels for machine learning datasets, it is important to assess the quality of crowd-provided data to improve analysis performance and reduce biases in subsequent machine learning tasks. Given the lack of ground truth in most cases of crowdsourcing, we refer to data quality as the annotators' consistency and credibility. Unlike the simple scenarios where kappa coefficient and intraclass correlation coefficient usually can apply, online crowdsourcing requires dealing with more complex situations. We introduce a systematic method for evaluating data quality and detecting spamming threats via variance decomposition, and we classify spammers into three categories based on their different behavioral patterns. A spammer index is proposed to assess entire data consistency, and two metrics are developed to measure crowd workers' credibility by utilizing the Markov chain and generalized random effects models. Furthermore, we demonstrate the practicality of our techniques and their advantages by applying them to a face verification task using both simulated and real-world data collected from two crowdsourcing platforms.
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
Data Collection by Survey
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Convenience Sampling Method
Convenience sampling is a non-random method of sample selection; this method selects individuals that are easily accessible and may result in biased data. For example, a marketing...
Detection of Gross Error: The Q Test
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...

