在隐藏混下对高维线性回归的解混和反混估计,并应用于omics数据
Zhaoyang Li1, Yahang Liu1, Kecheng Wei1
1Department of Biostatistics, School of Public Health, Fuda n University, Shanghai, 200032, China.
Bioinformatics (Oxford, England)
|July 14, 2025
概括
这项研究引入了一种新的两步方法,用于解决高维数据中隐藏的混问题,改善因果效应估计. 这种方法使用光谱转换和凸起式优化来进行准确的解混和脱,而无需事先了解混器.
科学领域:
- 统计 统计 统计 统计
- 生物信息学是一种生物信息学.
- 因果推理因果推理
背景情况:
- 观察性研究面临着高维数据中隐藏的混因素的挑战,导致偏见的因果效应估计.
- 现有的解混方法在高维设置中经常失败,或者需要先前了解混器.
研究的目的:
- 为隐藏混的高维线性回归提出一个两步解混和不混的估计方法.
- 开发一种不需要先前了解隐藏的混因素的技术.
主要方法:
- 一种涉及光谱转换的两步方法,用于消除混.
- 通过反转卡鲁什-库恩-塔克尔条件,使用凸优化进行偏差校正,而不假定精度矩阵稀疏.
主要成果:
- 提出的方法有效地减少了隐藏的混,并纠正了估计偏差.
- 与现有方法相比,模拟表明系数估计的精度提高了.
- 该方法应用于研究阿尔茨海默病严重程度和脑脊髓液水平的数据集.
结论:
- 这种新的解混和除技术为具有隐藏混因子的高维数据提供了强大的解决方案.
- 这种方法可以提高复杂数据集中的因果推理准确度.
相关概念视频
Confounding in Epidemiological Studies
258
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
258
Strategies for Assessing and Addressing Confounding
153
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
153
Bias in Epidemiological Studies
666
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
666
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
706
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
706
Confidence Interval for Estimating Population Mean
8.0K
A point estimate of the population mean is obtained from a single sample. Such a point estimate does not represent a population well because it needs to account for variability in the population. Single point estimate can also be biased despite the sample being selected randomly. Thus, a point estimate is often unreliable. A confidence interval is needed to reduce this unreliability.
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
8.0K
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K


