对于极高维度基因组数据的随机 LASSO
Beomsu Baek1, Jongkwon Jo2,3, Mingon Kang4
1Department of Computer Science, University of Nevada, Las Vegas, 89154, NV, USA.
Scientific reports
|January 14, 2026
概括
随机 LASSO 通过减少多对线性和采样随机性来增强高维基因组数据的特征选择. 这种新方法在识别显著生物标志物和估计系数方面表现优于现有模型.
科学领域:
- 基因组学就是基因组学.
- 生物统计学 生物统计学
- 机器学习 机器学习
背景情况:
- 高维数据分析对于基因组研究至关重要.
- 最小绝对收缩和选择操作员 (LASSO) 和其变体用于生物标志物发现.
- 现有的基于引导的LASSO模型面临着诸如多线性和采样随机性等挑战.
研究的目的:
- 介绍Stochastic LASSO,一种基于启动链的新方法来进行特征选择.
- 解决现有的LASSO模型在极高维度但低样本大小 (EHDLSS) 数据中的局限性.
- 改善特征选择准确性和基因组分析中的系数估计.
主要方法:
- 开发了基于引导的新方法Stochastic LASSO.
- 实施的技术,以减少多对线性和采样随机性.
- 使用了两阶段的t测试策略来确定统计学意义.
- 通过广泛的模拟和TCGA癌症基因表达数据来评估性能.
主要成果:
- 与基准模型相比,随机 LASSO 在特征选择和系数估计方面表现优异.
- 该方法有效地减少了多对线性和采样随机性.
- 在TCGA癌症数据集中确定了与生存预测相关的统计学上显著的基因.
- 在模拟实验中显示出更好的稳定性.
结论:
- 随机 LASSO 对 EHDLSS 基因组数据的现有 LASSO 模型提供了显著的进步.
- 拟议的方法为生物标志物发现和特征选择提供了强大而准确的工具.
- 随机 LASSO 在癌症基因组学中有实际应用,用于生存预测.
相关概念视频
Genome-wide Association Studies-GWAS
15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.3K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
1.1K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
1.1K
Quantifying and Rejecting Outliers: The Grubbs Test
3.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K
Single Nucleotide Polymorphisms-SNPs
17.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.9K


