对于一般化线性模型混合物的局部和整体偏差R平方尺度
Roberto Di Mari1, Salvatore Ingrassia1, Antonio Punzo1
1Dipartimento di Economia e Impresa, Università di Catania, Catania, Italy.
概括
这项研究引入了全局线性模型 (GLM) 混合物的新偏差测量方法,扩展R平方以更好地评估模型匹配. 这些措施用于分析COVID-19传播集群.
科学领域:
- 统计建模 统计建模
- 机器学习 机器学习
- 流行病学 流行病学
背景情况:
- 通用线性模型 (GLMs) 使用偏差来评估模型的合适性.
- 现有的R平方测量对于复杂的混合模型来说是有限的.
- 通过EM算法的最大概率 (ML) 是混合物模型参数估计的标准.
研究的目的:
- 扩大GLM混合物的偏差测量和R平方.
- 为集群层面和样本层面的分析制定本地和全球适应性措施.
- 将这些新措施应用于COVID-19传播数据.
主要方法:
- 将偏差测量扩展到使用ML和EM算法的GLM混合物.
- 提出局部和总偏差的正常化分解.
- 定义混合模型的局部和总偏差R平方尺度.
主要成果:
- 为GLM混合物开发了新的局部和全球偏差R平方测量方法.
- 通过对高斯式,波桑式和二项式响应的模拟来证明实用性.
- 应用措施分析意大利的COVID-19传播集群.
结论:
- 拟议的偏差R平方措施有效地评估GLM混合物中的模型合适性.
- 这些措施为集群分离和模型性能提供了可解释的见解.
- 该方法适用于现实世界流行病学数据分析.
相关概念视频
Variation
6.8K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.8K
Calibration Curves: Correlation Coefficient
1.7K
In a linear calibration curve, there is a value called the calibration coefficient, denoted by 'r,' which measures the strength and the direction of association between two variables. The correlation coefficient value ranges from −1 to +1. A value of +1 indicates a perfect positive linear correlation, −1 denotes a perfect negative correlation, and 0 implies no correlation between the two variables. A positive correlation value establishes that as one variable increases, the...
1.7K
Calculating and Interpreting the Linear Correlation Coefficient
6.0K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
6.0K
Coefficient of Correlation
6.2K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
6.2K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Friedman Two-way Analysis of Variance by Ranks
250
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
250


