在自然主义阅读中,单词频率和可预测性分离
1Department of Brain & Cognitive Sciences and McGovern Institute for Brain Research, Massachusetts Institute of Technology, Cambridge, MA, USA.
Open mind : discoveries in cognitive science
|March 13, 2024
概括
语言处理中的单词频率和可预测性效应是不同的认知现象. 这项研究证实了使用大规模数据和先进的计算模型在自然主义故事阅读中这些可分离的效应.
科学领域:
- 心理语言学 心理语言学
- 认知科学 认知科学
- 计算语言学 计算语言学
背景情况:
- 读者对不常见或不太可预测的单词的阅读时间较慢.
- 关于单词频率和可预测性效应是否是单独的认知过程存在争议.
- 关于分离的先前证据仅限于小样本,人工材料和简单的建模假设.
研究的目的:
- 调查词汇频率和可预测性效应是否在自然语言理解,特别是故事阅读过程中脱离.
- 通过使用大规模数据集和先进的计算建模来解决先前研究的局限性.
- 为了确定频率和可预测性效应在普通读取中是否可分离和添加.
主要方法:
- 对大量自然阅读数据 (六个数据集,超过220万个数据点) 的分析.
- 使用在广泛数据上训练的高级统计语言模型,估计单词频率和可预测性.
- 应用非线性连续时间回归来模型阅读行为.
主要成果:
- 这些发现支持单词频率和可预测性的可分离效应,与早期的实验研究一致.
- 发现频率和可预测性效应都是附加的.
- 该研究成功地使用自然主义数据和复杂的建模在规模上证明了这些分离.
结论:
- 单词频率和上下文可预测性代表了语言处理中的不同的认知现象.
- 这些可分离的效应甚至在自然主义的故事阅读中也可以观察到.
- 这些发现验证和扩展了以前的研究,使用更强大的方法和数据.
更多相关视频
相关概念视频
Language and Cognition
345
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
345
Group Design
8.9K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
8.9K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Determination of Expected Frequency
2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Naturalistic Observations
15.4K
If you want to understand how behavior occurs, one of the best ways to gain information is to simply observe the behavior in its natural context. However, people might change their behavior in unexpected ways if they know they are being observed. How do researchers obtain accurate information when people tend to hide their natural behavior? As an example, imagine that your professor asks everyone in your class to raise their hand if they always wash their hands after using the restroom. Chances...
15.4K
Predicting Products: Substitution vs. Elimination
11.6K
When a nucleophile and an alkyl halide react, nucleophilic substitution and β-elimination reactions compete to generate products.
The following factors can influence the mechanisms competing against each other:
The following factors can influence the mechanisms competing against each other:
11.6K


