高维的通用中位适应拉索与应用到omics数据.
Yahang Liu1, Qian Gao2,3, Kecheng Wei1
1Department of Biostatistics, School of Public Health, Fudan University, Shanghai, China.
Briefings in bioinformatics
|March 4, 2024
概括
通用中位适应拉索 (GMAL) 方法准确地选择变量并估计因果关系,即使结果数据偏差. 这种方法对高维数据集的现有方法进行了改进,特别是在现实世界的应用中.
科学领域:
- 统计 统计 统计 统计
- 生物统计学 生物统计学
- 基因组学就是基因组学.
背景情况:
- 高维数据分析在因果推理的变量选择中提出了挑战.
- 偏斜的结果分布可能会损害传统因果效应估计方法的准确性.
- 准确的共同变量选择对于可靠的因果推断至关重要.
研究的目的:
- 引入一种新的方法,即通用中位数自适应拉索 (GMAL),用于因果推理中的强大的变量选择.
- 为了应对高维数据中偏的结果分布带来的挑战.
- 为了提高因果效应估计在偏斜数据条件下的准确性.
主要方法:
- 开发了通用介质适应激光器 (GMAL) 用于共同变量选择.
- 使用线性中位数回归模型在GMAL中构建惩罚权重.
- 通过模拟和应用到现实世界数据集来评估GMAL的性能.
主要成果:
- GMAL显示了与对称结果分布的现有方法可比的变量选择性能.
- 当结果分布偏差时,GMAL在变量选择中表现优越.
- 在因果效应估计方面,GMAL的表现始终优于现有的方法,其证据是较低的平方根平均误差.
结论:
- GMAL有效地处理偏差结果分布,确保精确的变量选择和因果效应估计.
- 拟议的方法为高维设置中的因果推理提供了显著的进步,其结果非正常分布.
- 将GMAL应用于DNA甲基化数据集,以探索脑脊液tau蛋白与阿尔茨海默病严重程度之间的关系.
相关概念视频
Wilcoxon Signed-Ranks Test for Median of Single Population
129
The Wilcoxon signed-rank test for the median of a single population is a nonparametric test used to evaluate whether the median of a population differs from a specified value. Unlike parametric tests, it does not require data to follow a normal distribution, making it suitable for non-normal or small samples. The test begins by calculating the difference (d) between each observation and the hypothesized median. The absolute values of these differences are ranked in ascending order, with ties...
129
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
498
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
498
Genome-wide Association Studies-GWAS
13.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.4K
Kaplan-Meier Approach
137
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...
137
Friedman Two-way Analysis of Variance by Ranks
194
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
194
Median
18.5K
Besides mean, the median is a widely used measure of central tendency. Typically, median is defined as the central or middle value of a data set, measured by arranging the data elements in an increasing or decreasing order. Since this middle value is not affected by the precise numerical values of the outliers or fluctuations, it is insensitive to them. Hence, in cases where a data set may have outliers or the extreme values are not known, the median is a better measure of the central tendency...
18.5K


