由于未知数量的常见病例而导致的残余相关性,对双变数数据的关节疾病映射
Edouard Chatignoux1, Zoé Uhry1,2, Laurent Remontet2,3
1Sante publique France, Saint-Maurice 94415, France.
Biometrics
|September 1, 2025
概括
一个新的Bivariate-Poisson共享组件模型 (BP-SCM) 通过计算常见病例来准确分析相关疾病数量. 这改进了Poisson共享组件模型 (P-SCM),该模型在案例重叠时显示偏差.
科学领域:
- 生物统计学
- 空间流行病学
- 统计模型
背景情况:
- 双变数数据的联合空间分布通常使用Poisson共享组件模型 (P-SCM) 建模.
- P-SCM假设隐性变量完全解释了结果的相关性,当结果共享未知数量的情况时,这可能是不准确的.
- 这导致剩余相关性错误地归因于隐性变量,导致偏见的推断和糟糕的预测.
研究的目的:
- 解决P-SCM在模拟重叠病例的双变数数据中的局限性.
- 根据比瓦里亚特-波桑分布 (BP-SCM) 提出一种新的共享组件模型.
- 提高空间相关模型的准确性和联合计数结果的预测性能.
主要方法:
- 开发了一个共享组件模型 (BP-SCM),该模型将计数分解为共同和不同的情况.
- 使用高斯马科夫随机场建模了三种结果计数 (两个不同的,一个共同的).
- 采用贝叶斯框架与哈密尔顿蒙特卡洛推理进行模型估计.
主要成果:
- 拟议的BP-SCM与标准的P-SCM相比显示出更高的推断和预测性能.
- 模拟和现实应用证实了P-SCM在处理重叠案件时固有的偏差.
- BP-SCM成功估计了常见和不同病例的平均水平及其空间变化.
结论:
- BP-SCM为分析空间相关的双变量计数数据提供了更准确和更强大的方法,特别是当病例重叠时.
- 它克服了传统P-SCM的推断偏见和预测局限性.
- BP-SCM提供了关于共同和独特疾病模式及其空间依赖性的有价值的流行病学见解.
相关概念视频
Coefficient of Correlation
6.4K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
6.4K
Residuals and Least-Squares Property
7.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.8K
Correlation and Regression
1.8K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.8K
Correlation of Experimental Data
269
Dimensional analysis simplifies complex physical problems and guides experimental investigations, but it does not provide complete solutions. It identifies the dimensionless groups that influence a phenomenon, but experimental data is needed to establish the specific relationships and validate theoretical predictions.
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity,...
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity,...
269
Scatter Plot
8.7K
The most common and easiest way to display the relationship between two variables, x and y, is a scatter plot. A scatter plot shows the direction of a relationship between the variables. A clear direction happens when there is either:
8.7K
Calculating and Interpreting the Linear Correlation Coefficient
6.4K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
6.4K


