A large-scaled corpus for assessing text readability
Scott Crossley1, Aron Heintz2, Joon Suh Choi3
1Georgia State University, Atlanta, GA, USA. scrossley@gsu.edu.
Behavior Research Methods
|March 17, 2022
Summary
The CommonLit Ease of Readability (CLEAR) corpus offers unique readability scores for 5000 text excerpts. This resource aids researchers in developing and testing text readability metrics.
Area of Science:
- Computational Linguistics
- Educational Technology
- Psycholinguistics
Background:
- Readability metrics are crucial for understanding text complexity.
- Existing corpora may lack sufficient size, scope, or relevant metadata for advanced research.
- Developing robust readability models requires diverse and well-annotated text datasets.
Purpose of the Study:
- Introduce the CommonLit Ease of Readability (CLEAR) corpus.
- Provide researchers with a comprehensive resource for studying text readability and discourse processing.
- Facilitate the development and validation of new readability metrics.
Main Methods:
- Compilation of approximately 5000 text excerpts spanning over 250 years and multiple genres.
- Generation of unique readability scores based on teacher ratings of text difficulty.
- Inclusion of metadata such as publication year and genre for each excerpt.
Main Results:
- The CLEAR corpus contains ~5000 text excerpts with associated readability scores and metadata.
- The corpus covers a wide range of historical periods and genres, offering broad applicability.
- Reliability metrics for the human readability ratings have been established and are presented.
Conclusions:
- The CLEAR corpus represents a significant advancement over previous readability resources.
- It provides a valuable foundation for research in reading comprehension, discourse analysis, and educational technology.
- The corpus enables the development and testing of more accurate and nuanced readability assessment tools.
More Related Videos
Related Concept Videos
Goodness-of-Fit Test
4.4K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
4.4K
Protein Folding Quality Check in the RER
4.0K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
4.0K
Quantifying and Rejecting Outliers: The Grubbs Test
2.3K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.3K
Proofreading
6.8K
Synthesis of new DNA molecules is carried out by the enzyme DNA polymerase, which adds nucleotides on the daughter strand complementary to the template DNA strand. DNA polymerase has a higher affinity to add the correct base and ensures fidelity during DNA replication. Furthermore, it exhibits proofreading activity during replication, using an exonuclease domain that cuts off incorrect nucleotides from the nascent DNA strand.
Errors During Replication are Corrected by the DNA Polymerase...
Errors During Replication are Corrected by the DNA Polymerase...
6.8K
Sample Size Calculation
3.9K
Knowledge of the sample size is the first requirement to conduct random sampling or an experiment. The sample size is the total number of units, observations, or groups (in some cases) used to get the data to estimate a population parameter. As the name suggests, the sample size is that of the sample drawn from the population and differs from the population size.
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
3.9K
Lampbrush Chromosomes
2.5K
2.5K


