在读取过程中对目标回归进行信息理论分析
Ethan Gotlieb Wilcox1, Tiago Pimentel2, Clara Meister1
1Department of Computer Science, ETH Zürich, Switzerland.
Cognition
|May 21, 2024
概括
阅读过程中的向后顺序或回归被解释为信息理论方法. 关联的单词对,通过点向相互信息 (pmi) 测量,预测回归模式,支持重新激活假设.
科学领域:
- 认知科学 认知科学
- 计算语言学 计算语言学
- 心理语言学 心理语言学
背景情况:
- 在阅读过程中,回归 (反向的冲刺) 很常见 (占所有冲刺的5-20%).
- 回归的根本原因仍然不太清楚.
- 现有的假设 (重新激活,重新分析) 缺乏定量预测.
研究的目的:
- 使用信息理论来运行重新激活和重新分析假设.
- 预测基于词汇关联的回归模式 (点对点的相互信息,pmi).
- 通过使用预期的pmi (E[pmi]) 调查回归中的不确定性作用.
主要方法:
- 开发了一个信息理论框架来分析回归.
- 利用当代语言模型来估计单词对的 pmi 和 E[pmi].
- 分析了来自三个英语语体和六个非英语语体的语言家族的眼睛跟踪数据.
主要成果:
- 积极的pmi和E[pmi]值始终可以预测回归事件的发生.
- 负 pmi 和 E[pmi] 值不能预测回归模式.
- 结果在不同语言和语言模型中一致.
结论:
- 反激活假设得到了研究结果的支持.
- 一种信息理论方法提高了理解回归的预测能力.
- 本研究提供了阅读回归的第一个跨语言分析,并将其与语言处理中的信息理论原则联系起来.
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Correlation and Regression
1.2K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.2K
Censoring Survival Data
82
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
82
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Mechanistic Models: Compartment Models in Individual and Population Analysis
37
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
37


