Related Experiment Video
Updated: Jun 14, 2026

Experimental Paradigm for Measuring the Effect of Induced Emotion on Grammar Learning
Published on: January 29, 2020
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan1, Joanna J Bryson1,2, Arvind Narayanan1
1Center for Information Technology Policy, Princeton University, Princeton, NJ, USA. aylinc@princeton.edu jjb@alum.mit.edu arvindn@cs.princeton.edu.
Abstract:
Machine learning is a means to derive artificial intelligence by discovering patterns in existing data. Here, we show that applying machine learning to ordinary human language results in human-like semantic biases. We replicated a spectrum of known biases, as measured by the Implicit Association Test, using a widely used, purely statistical machine-learning model trained on a standard corpus of text from the World Wide Web. Our results indicate that text corpora contain recoverable and accurate imprints of our historic biases, whether morally neutral as toward insects or flowers, problematic as toward race or gender, or even simply veridical, reflecting the status quo distribution of gender with respect to careers or first names. Our methods hold promise for identifying and addressing sources of bias in culture, including technology.
Related Concept Videos
Stereotype Content Model
Natural and Artificial Concepts
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Automatic Processing and Automatic Social Behavior
