Related Experiment Video
Updated: Jan 12, 2026

08:32
Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
Published on: September 5, 2019
5.9K
An open dataset of Chinese duration expressions
Si-Qi Zhang1,2, Jia-Wen Niu1,2, Xiaoqian Liu1,2
1State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, Chinese Academy of Sciences, Beijing, China.
Scientific Data
|November 3, 2025
Summary
This study introduces a new dataset of 2,101 Chinese duration expressions, mapping verbal terms to numerical values. This resource aids research in natural language processing and linguistics by providing word frequencies for temporal expressions.
Area of Science:
- Linguistics
- Natural Language Processing
- Psychology
Background:
- Duration information is crucial for understanding text, presented numerically (e.g., 1 hour) or verbally (e.g., shortly).
- Existing lexicons lack comprehensive mappings between verbal duration expressions and numerical values.
- Databases of temporal expressions often omit word frequency data, hindering processing analysis.
Purpose of the Study:
- To create a comprehensive dataset of Chinese duration expressions with corresponding numerical values and word frequencies.
- To support research in natural language processing, psychology, and linguistics by providing a valuable resource for temporal information analysis.
- To address the gap in lexicons for converting verbal duration expressions to numerical durations.
Main Methods:
- Compiled an open dataset of 2,101 Chinese duration expressions.
- Annotated each expression with its corresponding numerical duration.
- Obtained word frequencies from a 10 billion character corpus (BLCU Corpus Center) and computed adjusted frequencies.
Main Results:
- A dataset of 2,101 Chinese duration expressions, each linked to a numerical duration.
- Inclusion of adjusted word frequencies for each expression, derived from a large-scale corpus.
- The dataset provides a foundational resource for analyzing temporal expressions in Chinese.
Conclusions:
- The newly created dataset offers a significant resource for researchers studying temporal information in Chinese.
- This dataset facilitates advancements in natural language processing, psychology, and linguistics by enabling quantitative analysis of duration expressions.
- The inclusion of word frequency data enhances the dataset's utility for understanding information processing related to time.
More Related Videos
Related Concept Videos
Interval Level of Measurement
17.9K
For effective statistical analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using the interval scale are similar to ordinal level data because they have a definite arrangement. However, in the interval level of measurement, the differences between data values are meaningful even though the data does not have a starting point.
Temperature is measured using the interval scale. It is measurable data, and the difference between...
Data measured using the interval scale are similar to ordinal level data because they have a definite arrangement. However, in the interval level of measurement, the differences between data values are meaningful even though the data does not have a starting point.
Temperature is measured using the interval scale. It is measurable data, and the difference between...
17.9K
Midrange
4.2K
A somewhat easy to compute quantitative estimate of a data set’s central tendency is its midrange, which is defined as the mean of the minimum and maximum values of an ordered data set.
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
4.2K

