一个通用线性混合效应模型,通过将多项研究中的追溯测量数据相结合,推断生物标志物相关性
Chengjie Xiong1,2, Ruijin Lu1,2, David Wolk3
1Division of Biostatistics, Washington University School of Medicine, St. Louis, Missouri, USA.
Statistics in medicine
|July 21, 2025
概括
生物医学研究可以通过协调平台将多项研究的生物标记数据结合起来. 一个新的元分析模型估计了真正的生物学相关性,提高了检测生物标志物和临床结果之间的关联的统计能力.
科学领域:
- 生物统计学 生物统计学
- 生物标志物发现发现
- 临床研究方法论 临床研究方法论
背景情况:
- 在生物医学研究中,大样本尺寸对于检测微妙的生物标志物-结果关联至关重要.
- 结合多项研究的数据增强了统计能力,但需要协调异构的生物标志物数据.
- 现有的方法在研究中与不同的平台和协议作斗争.
研究的目的:
- 开发一种新的元分析方法,用于协调跨研究的回顾性生物标志物数据.
- 估计潜伏生物标志物与临床结果之间的真实生物相关性.
- 为数据协调提供关于最佳桥接样本大小的指导.
主要方法:
- 使用测量误差模型概念化了一个潜在的生物标志物,用于观察到的版本.
- 开发了一种整体线性混合效应模型,集成相关性和类内相关系数 (ICC).
- 纳入随机效应以考虑研究异质性,并使用桥梁样本进行估计.
主要成果:
- 拟议的模型准确地估计了具有最小偏差 (≤0.03) 的生物相关性.
- 即使是小到平的ICC,也可以实现有效的协调.
- 对于大型ICC,只需要10%的桥梁样本才能进行公正的相关性估计.
结论:
- 超分析模型提供了一种可靠的方法来协调追溯生物标记数据.
- 它为确定必要的桥梁样本数量提供了有价值的指导.
- 该方法还可以在单个研究中解决批量效应.
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
87
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
87
Longitudinal Studies
248
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
248
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
Comparing the Survival Analysis of Two or More Groups
291
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
291
Bias in Epidemiological Studies
695
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
695
Longitudinal Research
12.5K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
12.5K


