一个开放的数据集中文持续时间表达式
Si-Qi Zhang1,2, Jia-Wen Niu1,2, Xiaoqian Liu1,2
1State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, Chinese Academy of Sciences, Beijing, China.
Scientific data
|November 3, 2025
概括
这项研究引入了2,101个中文持续时间表达式的新数据集,将口头术语映射到数值. 该资源通过提供时间表达式的单词频率来帮助自然语言处理和语言学的研究.
科学领域:
- 语言学的语言学.
- 自然语言处理自然语言处理.
- 心理学 心理学 心理学
背景情况:
- 持续时间信息对于理解文本至关重要,以数字形式 (例如,1小时) 或口头形式 (例如,很快) 呈现.
- 现有的词典缺乏语言持续时间表达式和数字值之间的全面映射.
- 时间表达式的数据库经常省略单词频率数据,阻碍处理分析.
研究的目的:
- 创建一个全面的数据集的中国时间表达式与相应的数值和单词频率.
- 通过为时间信息分析提供宝贵的资源,支持自然语言处理,心理学和语言学方面的研究.
- 为了弥补词典中的差距,用于将口头持续时间表达式转换为数值持续时间.
主要方法:
- 编制了 2,101 个中文持续时间表达式的开放数据集.
- 标注每个表达式及其相应的数值持续时间.
- 从100亿个字符库 (BLCU Corpus Center) 中获得单词频率,并计算调整频率.
主要成果:
- 一个包含 2,101 个中文持续时间表达式的数据集,每个表达式都与一个数字持续时间相关联.
- 包括每个表达式的调整后的单词频率,从大规模的语料库中获得.
- 该数据集为分析中文时间表达式提供了基础资源.
结论:
- 新创建的数据集为研究人员研究中文时间信息提供了重要的资源.
- 这一数据集通过对持续时间表达式进行定量分析,促进了自然语言处理,心理学和语言学的进步.
- 包含单词频率数据提高了数据集的实用性,以了解与时间相关的信息处理.
相关概念视频
Interval Level of Measurement
17.9K
For effective statistical analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using the interval scale are similar to ordinal level data because they have a definite arrangement. However, in the interval level of measurement, the differences between data values are meaningful even though the data does not have a starting point.
Temperature is measured using the interval scale. It is measurable data, and the difference between...
Data measured using the interval scale are similar to ordinal level data because they have a definite arrangement. However, in the interval level of measurement, the differences between data values are meaningful even though the data does not have a starting point.
Temperature is measured using the interval scale. It is measurable data, and the difference between...
17.9K
Midrange
4.2K
A somewhat easy to compute quantitative estimate of a data set’s central tendency is its midrange, which is defined as the mean of the minimum and maximum values of an ordered data set.
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
4.2K


