Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Determination of Expected Frequency01:08

Determination of Expected Frequency

2.3K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.3K
Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

2.7K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
2.7K
Relative Frequency Distribution00:55

Relative Frequency Distribution

11.7K
A relative frequency distribution is the proportion or fraction of times a value occurs in a data set. To find the relative frequencies, one can divide each frequency by the total number of data points in the sample. It is very similar to a regular frequency distribution, except that instead of reporting how many data values fall in a class, a relative frequency distribution reports the fraction of data values that fall in a class. These fractions or proportions are called relative frequencies...
11.7K
Prediction Intervals01:03

Prediction Intervals

2.4K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
2.4K
Wald-Wolfowitz Runs Test I01:17

Wald-Wolfowitz Runs Test I

749
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
749
Frequency-dependent Selection01:21

Frequency-dependent Selection

22.3K
When the fitness of a trait is influenced by how common it is (i.e., its frequency) relative to different traits within a population, this is referred to as frequency-dependent selection. Frequency-dependent selection may occur between species or within a single species. This type of selection can either be positive—with more common phenotypes having higher fitness—or negative, with rarer phenotypes conferring increased fitness.
22.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Statistical errors undermine claims about the evolution of polysynthetic languages.

Proceedings of the National Academy of Sciences of the United States of America·2025
Same author

Still No Evidence for an Effect of the Proportion of Non-Native Speakers on Natural Language Complexity.

Entropy (Basel, Switzerland)·2024
Same author

Languages with more speakers tend to be harder to (machine-)learn.

Scientific reports·2023
Same author

A large quantitative analysis of written language challenges the idea that all languages are equally complex.

Scientific reports·2023
Same author

Is More Always Better? Testing the Addition Bias for German Language Statistics.

Cognitive science·2023
Same author

Studying Lexical Dynamics and Language Change via Generalized Entropies: The Problem of Sample Size.

Entropy (Basel, Switzerland)·2020

Related Experiment Video

Updated: Sep 21, 2025

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.3K

Testing the Relationship between Word Length, Frequency, and Predictability Based on the German Reference Corpus.

Alexander Koplenig1, Marc Kupietz2, Sascha Wolfer1

  • 1Department of Lexical Studies, Leibniz-Institute for the German Language (IDS).

Cognitive Science
|June 6, 2022
PubMed
Summary

This study challenges the idea that information content predicts word length better than frequency. Analyzing a large German corpus, we found minimal support for this linguistic theory.

Keywords:
CompressionCorpus linguisticsInformation theoryLarge-scale corporaN-gram modelingUniform information density

More Related Videos

Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese
08:08

Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese

Published on: April 1, 2016

9.4K
Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

922

Related Experiment Videos

Last Updated: Sep 21, 2025

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.3K
Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese
08:08

Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese

Published on: April 1, 2016

9.4K
Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

922

Area of Science:

  • Computational Linguistics
  • Corpus Linguistics
  • Psycholinguistics

Background:

  • Investigates the relationship between word length, frequency, and information content.
  • Revisits the findings of Piantadosi, Tily, and Gibson (2011) on word length prediction.
  • Addresses methodological challenges in analyzing large-scale linguistic corpora, as highlighted by Meylan and Griffiths (2021).

Discussion:

  • The study tested the hypothesis that average information content is a superior predictor of word length compared to word frequency.
  • Analysis was conducted on a substantial subset of the German Reference Corpus, comprising approximately 43 billion words.
  • Results indicate limited empirical support for the central claim made by Piantadosi, Tily, and Gibson (2011).

Key Insights:

  • Word frequency appears to be a more robust predictor of word length than information content in the analyzed German corpus.
  • The findings suggest that the proposed information-theoretic explanation for word length may not generalize across all languages or corpora.
  • Methodological considerations in corpus linguistics are crucial for validating theoretical claims.

Outlook:

  • Further research is needed to explore the predictive power of information content across diverse languages and text types.
  • Investigating alternative or complementary factors influencing word length in linguistic corpora is warranted.
  • Refining methodologies for corpus analysis may yield clearer insights into the relationship between linguistic properties.