Related Experiment Video
Updated: Apr 1, 2026

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Characterizing the Google Books Corpus: Strong Limits to Inferences of Socio-Cultural and Linguistic Evolution
Eitan Adam Pechenick1, Christopher M Danforth1, Peter Sheridan Dodds1
1Department of Mathematics and Statistics, University of Vermont, Burlington, Vermont, United States of America; Center for Complex Systems, University of Vermont, Burlington, Vermont, United States of America; Computational Story Lab, University of Vermont, Burlington, Vermont, United States of America; Vermont Advanced Computing Core, University of Vermont, Burlington, Vermont, United States of America.
Abstract:
It is tempting to treat frequency trends from the Google Books data sets as indicators of the "true" popularity of various words and phrases. Doing so allows us to draw quantitatively strong conclusions about the evolution of cultural perception of a given topic, such as time or gender. However, the Google Books corpus suffers from a number of limitations which make it an obscure mask of cultural popularity. A primary issue is that the corpus is in effect a library, containing one of each book. A single, prolific author is thereby able to noticeably insert new phrases into the Google Books lexicon, whether the author is widely read or not. With this understood, the Google Books corpus remains an important data set to be considered more lexicon-like than text-like. Here, we show that a distinct problematic feature arises from the inclusion of scientific texts, which have become an increasingly substantive portion of the corpus throughout the 1900 s. The result is a surge of phrases typical to academic articles but less common in general, such as references to time in the form of citations. We use information theoretic methods to highlight these dynamics by examining and comparing major contributions via a divergence measure of English data sets between decades in the period 1800-2000. We find that only the English Fiction data set from the second version of the corpus is not heavily affected by professional texts. Overall, our findings call into question the vast majority of existing claims drawn from the Google Books corpus, and point to the need to fully characterize the dynamics of the corpus before using these data sets to draw broad conclusions about cultural and linguistic evolution.
More Related Videos
06:33Decomposing the Variance in Reading Comprehension to Reveal the Unique and Common Effects of Language and Decoding
Published on: October 11, 2018
07:34Perceptual and Category Processing of the Uncanny Valley Hypothesis' Dimension of Human Likeness: Some Methodological Issues
Published on: June 3, 2013
Related Concept Videos
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Components of Language
Criticisms of the Evolutionary Perspective
Evolutionary psychology provides one explanation for these findings, suggesting...
Limits to Natural Selection
Complementation Tests
Organisms heterozygous for different mutations are crossed pairwise in all combinations. If present on different genes, the mutations can complement each other by providing the missing...