一个基于模型的集群算法与共变量调整及其应用到肺癌分层
Carlos E M Relvas1, Asuka Nakata2, Guoan Chen3
1Institute of Mathematics and Statistics, University of São Paulo, Rua do Matão 1010 São Paulo, São Paulo 05508-090, Brazil.
Journal of bioinformatics and computational biology
|September 11, 2023
概括
我们开发了CEM-Co,这是一种新的集群算法,可以最大限度地减少协变效应. 这种方法在肺癌患者中发现了一个预后较差的亚组,其表现优于标准算法.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 聚类对于数据分析至关重要,揭示隐藏的模式并产生假设.
- 共变量可以掩盖实证数据中的真实集群结构,导致误解.
- 标准聚类可能会根据混因素 (例如年龄) 而不是疾病状况将受试者分组.
研究的目的:
- 引入CEM-Co,一种基于模型的聚类算法,旨在减轻不良共变量的影响.
- 为了证明CEM-Co在复杂的生物数据集中识别有意义的子组的有效性.
主要方法:
- 开发了CEM-Co,一种基于模型的新型集群算法.
- 将CEM-Co应用于来自129名I期非小细胞肺癌患者的基因表达数据集.
- 将CEM-Co的性能与标准集群算法进行比较.
主要成果:
- CEM-Co成功地确定了一组肺癌患者的亚组,这些患者的预后较差.
- 标准集群算法未能检测出这种临床相关的子组.
- 在集群过程中,CEM-Co有效地消除或最大限度地减少了共变量的混效应.
结论:
- 在生物数据分析中,CEM-Co提供了一种可靠的方法,用于对共变量进行调整的聚类.
- 这种方法提高了发现临床显著子组的能力,改善了预后识别.
- 在处理共变量影响数据的各个领域,CEM-Co具有潜在的应用.
相关概念视频
Statistical Methods for Analyzing Epidemiological Data
403
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
403
Cancer Survival Analysis
381
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
381
Comparing the Survival Analysis of Two or More Groups
220
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
220
Survival Tree
105
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
105
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Strategies for Assessing and Addressing Confounding
119
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
119


