探索回归稀释偏差使用重复测量2858个变量的≤49 000英国生物库参与者
Charlotte E Rutter1,2, Louise A C Millard1,3, Maria Carolina Borges1,3
1MRC Integrative Epidemiology Unit, University of Bristol, Bristol, UK.
International journal of epidemiology
|June 19, 2023
概括
英国生物库数据中的测量错误很常见,可能会导致结果偏差. 纠正这种随机错误加强了暴露结果的关联,特别是红细胞分布宽度 (RDW) 和C反应蛋白 (CRP).
科学领域:
- 流行病学 流行病学
- 生物统计学 生物统计学
- 生物信息学是一种生物信息学.
背景情况:
- 暴露和混因子的测量错误是流行病学研究中的一个重大挑战.
- 这种错误可能导致暴露和健康结果之间的偏见关联.
- 在分析中,这种偏见往往被忽视,特别是在像英国生物银行这样的大型生物银行中.
研究的目的:
- 系统地评估在英国生物银行中的众多连续变量中的随机测量误差.
- 调查减轻测量误差对暴露结果关联的影响的方法.
- 为研究人员提供有价值的统计数据和工具,以解决测量错误.
主要方法:
- 计算了所有连续变量的类内相关系数 (ICC),并重复测量以量化随机误差.
- 使用回归校准来纠正发现的随机测量错误.
- 红细胞分布宽度 (RDW),C反应蛋白 (CRP) 和25-氧维生素D [25(OH) D]与死亡率之间的关联被用作案例研究.
主要成果:
- 总共评估了2858个连续变量,ICC在数据类型之间差异很大 (例如,成像测量:0.85,饮食测量:0.35).
- 具体的RDW,CRP和25(OH) D的ICC分别为0.52,0.29和0.55,表明错误不可以忽略.
- 对暴露的测量误差进行校正,加强了观察到的与死亡率的关联 (例如,RDW,CRP,25...OH) D).
结论:
- 随机测量错误很普遍,在大型生物库数据集中通常是相当大的.
- 这些发现强调了考虑和纠正测量误差的重要性,以获得更准确的暴露-结果关联.
- 该研究为研究界提供了必要的数据和方法,以解决他们自己的分析中的测量错误.
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Bias in Epidemiological Studies
380
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
380
Confounding in Epidemiological Studies
198
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
198
Longitudinal Studies
191
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
191
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Mechanistic Models: Compartment Models in Individual and Population Analysis
66
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
66


