一个非负的尖和板拉索概括的线性堆叠预测建模方法,用于高维的奥米克数据
Junjie Shen1, Shuo Wang2, Yongfei Dong1
1Department of Biostatistics, School of Public Health, Jiangsu Key Laboratory of Preventive and Translational Medicine for Geriatric Diseases, MOE Key Laboratory of Geriatric Diseases and Immunology, Suzhou Medical College of Soochow University, No. 199 Renai Road, Suzhou, 215123, Jiangsu, People's Republic of China.
BMC bioinformatics
|March 21, 2024
概括
这项研究引入了一种新的堆叠方法,使用非负的尖和板拉索 (nsslasso) 用于用omics数据预测疾病风险. 纳斯拉索堆叠方法提高了预测准确性,并确定了关键的生物结构.
科学领域:
- 基因组学就是基因组学.
- 生物统计学 生物统计学
- 计算生物学 计算生物学
背景情况:
- 高维的奥米克数据对于疾病风险预测至关重要.
- 现有的稀疏方法通常依赖于单个模型,导致局限性.
- 之前的生物知识可以指导OMIC数据分析.
研究的目的:
- 开发一种新的堆叠策略,用于使用高维的OMICS数据进行疾病风险预测.
- 与单一模型方法相比,提高预测准确性和概括性.
- 为了有效地利用生物组结构信息.
主要方法:
- 提出了一个使用非负的尖和板拉索 (nsslasso) 通用线性模型 (GLM) 的堆叠策略.
- 将omics数据细分成基于生物知识的子数据.
- 在每个细分上训练了子模型,并使用nsslasso GLM的超级学习器进行组合预测.
主要成果:
- 在模拟和真实世界乳腺癌数据上,nsslasso堆叠方法在单一模型方法上表现出优异的预测性能.
- 在处理冗余和识别重要的子模型方面,超越了传统的堆叠方法.
- 在高噪音环境中实现了强大的优异预测.
结论:
- 恩斯拉索方法提供了增强的预测准确性,稳定性和生物解释性.
- 它有效地识别了关键的生物组结构和潜在的新生物标志物.
- 这种方法促进了omics数据在临床和公共卫生研究中的应用.
相关概念视频
End Point Prediction: Gran Plot
322
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
322
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
491
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
491
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
69
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
69
Improving Translational Accuracy
10.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.3K
Genome-wide Association Studies-GWAS
13.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.4K


