识别隐性人口特征对风险因素的影响修改,具有稀疏变化的回归系数
bioRxiv : the preprint server for biology
|December 16, 2024
概括
这项研究引入了一种新的回归模型,该模型解释了风险因素与疾病结果的关联如何根据隐藏的个体特征而变化. 该方法改善了疾病风险预测,并揭示了肺癌数据中重要的效果修改.
科学领域:
- 流行病学 流行病学
- 生物统计学 生物统计学
- 机器学习 机器学习
- 基因组学就是基因组学.
背景情况:
- 观察数据分析对于了解疾病风险因素和结果至关重要.
- 传统模型经常假设固定的关联,可能缺少复杂的个体变化.
- 影响这些关联的潜在特征可能被观察到的风险因素部分捕获.
研究的目的:
- 开发一种新的回归模型,将潜在特征的效果修改纳入其中.
- 通过捕捉动态关联来提高疾病风险预测的准确性.
- 在现实数据集中识别潜在效应修改,例如肺癌.
主要方法:
- 开发了一种新的回归模型,其中系数因潜在特征的函数而变化.
- 从观察到的风险因素数据中提取隐藏特征.
- 使用模拟研究验证了该模型,并将其应用于癌症基因组图谱 (TCGA) 肺癌数据集.
主要成果:
- 拟议的模型在模拟中在各种数据设置中展示了卓越的性能.
- 应用到肺癌数据显示,与拉索和弹性网相比,预测准确度显著提高 (6%-118%AUC-0.5改善).
- 确定了与肺癌中特定基因通路相关的新型潜伏效应修饰.
结论:
- 新的回归方法有效地通过建模不同的关联来捕捉潜在的数据结构.
- 这种方法增强了疾病风险预测,并提供了对疾病病因学的更深入的了解.
- 这些发现强调了在流行病学研究中考虑潜在效应修改的重要性.
相关概念视频
Statistical Methods for Analyzing Epidemiological Data
299
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
299
Relative Risk
116
Relative risk (RR) is a statistical measure commonly used in epidemiology to compare the likelihood of a particular event occurring between two groups. This metric is important for evaluating the relationship between exposure to a specific risk factor and the probability of a particular outcome. It plays a crucial role in medical research, public health studies, and risk assessment. Relative risk quantifies how much more (or less) likely an event is to occur in an exposed group compared to an...
116
Mechanistic Models: Compartment Models in Individual and Population Analysis
27
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
27
Parametric Survival Analysis: Weibull and Exponential Methods
356
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
356
Strategies for Assessing and Addressing Confounding
83
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
83
Confounding in Epidemiological Studies
143
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
143


