Zipf's Law for Word Frequencies: Word Forms versus Lemmas in Long Texts

Álvaro Corral1, Gemma Boleda2, Ramon Ferrer-i-Cancho3

  • 1Centre de Recerca Matemàtica, Bellaterra, Barcelona, Spain; Departament de Matemàtiques, Universitat Autònoma de Barcelona, Bellaterra, Barcelona, Spain.

Plos One
|July 10, 2015
PubMed
Summary

Zipf's law applies to both word forms and lemma forms across languages. While exponents remain similar, the transformation from words to lemmas significantly impacts low-frequency cut-offs, causing them to increase.

Related Concept Videos

Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
8.9K
Poisson Probability Distribution01:09

Poisson Probability Distribution

A Poisson probability distribution is a discrete probability distribution. It gives the probability of a number of events occurring in a fixed interval of time or space if these events happen at a known average rate and independently of the time since the last event. For example, a book editor might be interested in the number of words spelled incorrectly in a particular book. It might be that, on average, there are five words spelled incorrectly in 100 pages. The interval is 100 pages.
The...
12.5K
Cumulative Frequency Distribution01:04

Cumulative Frequency Distribution

A cumulative frequency distribution is another type of frequency distribution. Instead of reporting how many data values fall in some classes, it reports how many data values are contained in either that class or any class to its left. Technically, it means the sum of frequencies of the class and all the classes below it in a frequency distribution. A cumulative frequency is calculated by adding the frequency of each class lower than the corresponding class interval or category. In general, a...
9.0K
Frequency-dependent Selection01:21

Frequency-dependent Selection

When the fitness of a trait is influenced by how common it is (i.e., its frequency) relative to different traits within a population, this is referred to as frequency-dependent selection. Frequency-dependent selection may occur between species or within a single species. This type of selection can either be positive—with more common phenotypes having higher fitness—or negative, with rarer phenotypes conferring increased fitness.
24.5K
Relative Frequency Distribution00:55

Relative Frequency Distribution

A relative frequency distribution is the proportion or fraction of times a value occurs in a data set. To find the relative frequencies, one can divide each frequency by the total number of data points in the sample. It is very similar to a regular frequency distribution, except that instead of reporting how many data values fall in a class, a relative frequency distribution reports the fraction of data values that fall in a class. These fractions or proportions are called relative frequencies...
14.3K
Percentage Frequency Distribution00:57

Percentage Frequency Distribution

A percentage frequency distribution, in general, is a display of data that indicates the percentage of observations for each data point or grouping of data points. It is a commonly used method for expressing the relative frequency of survey responses and other data. The percentage frequency distributions are often displayed as bar graphs, pie charts, or tables.
The process of making a percentage frequency distribution involves the following few steps: note the total number of observations;...
65.2K