Related Experiment Video
Updated: Sep 9, 2025

Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
Published on: May 9, 2017
The Indo-European Cognate Relationships dataset
Cormac Anderson1,2, Matthew Scarborough3, Lechosław Jocz4
1Department of Linguistic and Cultural Evolution, Max Planck Institute for Evolutionary Anthropology, Deutscher Platz 6, 04103, Leipzig, Germany. cormacanderson@gmail.com.
The Indo-European Cognate Relationships (IE-CoR) dataset maps word relationships across 160 Indo-European languages. This benchmark resource aids computational research into language evolution and historical linguistics.
Area of Science:
- Computational Linguistics
- Historical Linguistics
- Language Evolution Studies
Background:
- The Indo-European language family is vast, with complex historical relationships.
- Understanding cognate patterns is crucial for reconstructing proto-languages and tracing linguistic evolution.
- Existing datasets often lack comprehensive coverage or standardized formats for computational analysis.
Purpose of the Study:
- To introduce the Indo-European Cognate Relationships (IE-CoR) dataset, a novel resource for linguistic research.
- To provide a benchmark dataset for computational studies on Indo-European language evolution.
- To facilitate the study of lexical inheritance and horizontal transfer across Indo-European languages.
Main Methods:
- Compilation of a relational dataset covering 160 Indo-European languages.
- Analysis of 25,731 lexeme entries into 4,981 cognate sets based on 170 reference meanings.
- Inclusion of data on horizontal transfer, time calibration, and geographical/social metadata.
- Adherence to Cross-Linguistic Data Format (CLDF) protocols for interoperability.
Main Results:
- The IE-CoR dataset contains 25,731 lexeme entries across 160 languages, organized into 4,981 cognate sets.
- It represents all 13 main Indo-European clades and includes data on horizontal word transfer.
- The dataset incorporates time calibration, geographical, and social metadata for enhanced analysis.
Conclusions:
- The IE-CoR dataset offers a comprehensive, standardized resource for computational historical linguistics.
- It serves as a valuable benchmark for research into Indo-European language evolution.
- The dataset's design promotes interoperability and can serve as a model for other language families.
More Related Videos
07:28Studying Metabolic Brain Connectivity Using 2-Deoxy-2-[18F]Fluoro-D-Glucose Dynamic Positron Emission Tomography at the Single-subject Level
Published on: January 24, 2025
12:49Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Kendall's Coefficient of Concordance
Synteny and Evolution
Around 80 million years ago, the human and mice lineages diverged from the common ancestor. During the course of evolution, the ancestral...
Language and Cognition
Phylogeny
Convergent Evolution