Related Experiment Video
Updated: Jun 12, 2025

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Language Statistics at Different Spatial, Temporal, and Grammatical Scales
Fernanda Sánchez-Puig1,2,3, Rogelio Lozano-Aranda1,2, Dante Pérez-Méndez2,4
1Facultad de Ciencias, Universidad Nacional Autónoma de México, Mexico City 04510, Mexico.
This study analyzed English and Spanish on Twitter, finding word patterns (ngrams) vary most with grammar complexity. Rank diversity shows universal trends but national and temporal factors influence language use.
Area of Science:
- Computational Linguistics
- Quantitative Linguistics
- Sociolinguistics
Background:
- The availability of large datasets has propelled statistical linguistics.
- Twitter data offers a rich resource for analyzing real-time language use.
Purpose of the Study:
- To investigate rank diversity in English and Spanish using Twitter data.
- To examine the influence of temporal, spatial, and grammatical scales on language variation.
- To quantify universal language statistics and identify sources of variation.
Main Methods:
- Analysis of Twitter data from 2014 across eight countries.
- Investigation of word ngrams (1-grams to 5-grams) across temporal (3-96h) and spatial (3km-3000km) scales.
- Examination of rank diversity curves and statistical properties of Twitter-specific tokens (emojis, hashtags, mentions).
Main Results:
- Rank diversity shows similarity at the 1-gram level across languages, countries, and scales.
- Increased grammatical complexity (higher ngrams) leads to more pronounced variations influenced by temporal, spatial, linguistic, and national factors.
- Twitter-specific tokens exhibit a sigmoid pattern in their rank diversity function.
Conclusions:
- Grammatical scale is a key driver of language rank diversity variation.
- While universal patterns exist, language use on Twitter is shaped by context (time, location, nationality).
- The study quantifies language statistics and highlights factors influencing linguistic diversity in digital communication.
More Related Videos
05:15The Spatial Memory Game: Testing the Relationship Between Spatial Language, Object Knowledge, and Spatial Cognition
Published on: February 19, 2018
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Ordinal Level of Measurement
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
Nominal Level of Measurement
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal...
Interval Level of Measurement
Data measured using the interval scale are similar to ordinal level data because they have a definite arrangement. However, in the interval level of measurement, the differences between data values are meaningful even though the data does not have a starting point.
Temperature is measured using the interval scale. It is measurable data, and the difference between...
Levels of Use of a GIS
Selected Data About Geographic Locations
Components of Language