一个规范化的考克斯等级模型,用于将注释信息纳入预测性欧米研究中的信息
Dixin Shen1, Juan Pablo Lewinger2, Eric Kawaguchi2
1Clinical Data Science, Gilead Sciences, Foster City, USA. dixinshen@gmail.com.
BioData mining
|October 25, 2024
概括
将外部元特征与新的规范化层次框架集成,显著提高了时间到事件结果的预测准确性. 这种方法改善了高维的奥米克数据中的特征选择和发现,即使在信息不丰富的元特征中也提供了强大的性能.
科学领域:
- 生物信息学是一种生物信息学.
- 统计基因组学 统计基因组学
- 计算生物学 计算生物学
背景情况:
- 高维的奥米克数据通常包括诸如生物路径和功能注释之类的信息性元特征.
- 这些元特征可以增强对结果的预测,特别是时间到事件数据.
研究的目的:
- 引入一个规范化的层次框架,用于将元特征与omics数据集成.
- 为了改善预测和特征选择性能,以获得时间到事件的结果.
主要方法:
- 开发了一个层次框架,以纳入元特征.
- 在omics和meta-feature层面应用规范化来处理高维数据.
- 该模型是使用代重权最小方形和循环坐标下降来装配的.
主要成果:
- 与标准的考克斯回归相比,当元特征具有信息性时,规范化的层次模型大大提高了预测性能.
- 对乳腺癌和黑色素瘤存活率数据的应用证明了更好的预测,并确定了重要的omics特征集.
- 该模型显示出稳健性,在元特征信息不丰富时,其性能与标准方法相比较.
结论:
- 层次调节回归模型有效地整合了外部元特征信息,以获得时间到事件的结果.
- 它提高了预测准确度,并有助于发现重要特征,为预测和发现应用提供服务.
- 该框架是强大的,即使在没有信息的外部数据上也保持了性能.
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
29
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
29
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
399
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
399
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
96
Drug disposition in the body is a complex process and can be studied using two major approaches: the model and the model-independent approaches.
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
96
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
60
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
60
Statistical Methods for Analyzing Epidemiological Data
310
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
310


