Related Experiment Video
Updated: Oct 3, 2025

Design and Analysis for Fall Detection System Simplification
Published on: April 6, 2020
Machine learning to detect invalid text responses: Validation and comparison to existing detection methods
Ryan C Yeung1, Myra A Fernandes2
1Department of Psychology, University of Waterloo, Psychology, Anthropology, and Sociology (PAS) Building, 200 University Avenue West, Waterloo, ON, N2L 3G1, Canada. rcyeung@uwaterloo.ca.
Abstract:
A crucial step in analysing text data is the detection and removal of invalid texts (e.g., texts with meaningless or irrelevant content). To date, research topics that rely heavily on analysis of text data, such as autobiographical memory, have lacked methods of detecting invalid texts that are both effective and practical. Although researchers have suggested many data quality indicators that might identify invalid responses (e.g., response time, character/word count), few of these methods have been empirically validated with text responses. In the current study, we propose and implement a supervised machine learning approach that can mimic the accuracy of human coding, but without the need to hand-code entire text datasets. Our approach (a) trains, validates, and tests on a subset of texts manually labelled as valid or invalid, (b) calculates performance metrics to help select the best model, and (c) predicts whether unlabelled texts are valid or invalid based on the text alone. Model validation and evaluation using autobiographical memory texts indicated that machine learning accurately detected invalid texts with performance near human coding, significantly outperforming existing data quality indicators. Our openly available code and instructions enable new methods of improving data quality for researchers using text as data.
Related Concept Videos
Data Validation
Key parameters for method validation include:
Detection of Gross Error: The Q Test
Types of Errors: Detection and Minimization
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Quantifying and Rejecting Outliers: The Grubbs Test
Response Surface Methodology
The process of RSM involves several key steps:

