Related Experiment Video
Updated: Jun 15, 2026

08:04
Conditions Affecting Social Space in Drosophila melanogaster
Published on: November 5, 2015
Random texts do not exhibit the real Zipf's law-like rank distribution
Ramon Ferrer-I-Cancho1, Brita Elvevåg
1Departament de Llenguatges i Sistemes Informàtics, Universitat Politècnica de Catalunya, Barcelona, Catalonia, Spain. rferrericancho@lsi.upc.edu
Plos One
|March 17, 2010
Summary
Zipf's law, a linguistic principle, appears to hold true for natural languages. Statistical tests reveal that random texts do not accurately mimic the word rank distributions found in real texts.
Area of Science:
- Computational Linguistics
- Statistical Analysis
- Natural Language Processing
Background:
- Zipf's law describes the relationship between word frequency and rank in texts.
- Previous studies suggested random texts also follow Zipf's law, questioning its significance for language.
Purpose of the Study:
- To rigorously examine the validity of Zipf's law in random text generation.
- To determine if random texts truly exhibit the same rank distributions as natural languages.
Main Methods:
- Employed three distinct statistical tests to compare rank distributions.
- Analyzed both simple random texts and more complex, realistic random text models.
- Tested consistency even when parameters were derived from real text data.
Main Results:
- Statistical tests demonstrated significant inconsistencies between random and real text rank distributions.
- The purported "good fit" of random texts to Zipf's law was found to be flawed.
- Findings held true across various random text generation methods.
Conclusions:
- The claim that random texts accurately replicate Zipf's law distributions is unsubstantiated.
- Zipf's law may indeed be a fundamental characteristic of natural language structure.
- This research reinforces the linguistic significance of Zipf's law.
Related Concept Videos
Wald-Wolfowitz Runs Test II
The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
Probability Distributions
The probability of a random variable x is the likelihood of its occurrence. A probability distribution represents the probabilities of a random variable using a formula, graph, or table. There are two types of probability distribution– discrete probability distribution and continuous probability distribution.
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson probability...
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson probability...
Uniform Distribution
The uniform distribution is a continuous probability distribution of events with an equal probability of occurrence. This distribution is rectangular.Two essential properties of this distribution are The area under the rectangular shape equals 1. There is a correspondence between the probability of an event and the area under the curve.Further, the mean and standard deviation of the uniform distribution can be calculated when the lower and upper cut-offs, denoted as a and b,...
Distribution of Molecular Speeds
The motion of molecules in a gas is random in magnitude and direction for individual molecules, but a gas of many molecules has a predictable distribution of molecular speeds. This predictable distribution of molecular speeds is known as the Maxwell-Boltzmann distribution. The distribution of molecular speeds in liquids is comparable to that of gases but not identical and can help to understand the phenomenon of the boiling and vapor pressure of a liquid. Consider that a molecule requires a...
z Scores and Unusual Values
The z score is one of the three measures of relative standing. It describes the location of a value in a dataset relative to the mean. z scores are obtained after the standardization of the values in a dataset. The z score for the mean is 0.
This score indicates how far a value is from the mean in terms of standard deviation. For example, if a data value has a z score of +1, the researcher can infer that the particular data value is one standard deviation above the mean. If another data value...
This score indicates how far a value is from the mean in terms of standard deviation. For example, if a data value has a z score of +1, the researcher can infer that the particular data value is one standard deviation above the mean. If another data value...
Wald-Wolfowitz Runs Test I
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
The test works...
