联合聚类多个纵向特征:方法和软件包与实用指南的比较
Zihang Lu1,2, Mojtaba Ahmadiankalati1, Zhiwen Tan1
1Department of Public Health Sciences, Queen's University, Kingston, Ontario, Canada.
Statistics in medicine
|October 4, 2023
概括
这项研究指导研究人员在医学数据中聚类多个纵向特征,以发现疾病轨迹. 它比较了基于模型和基于算法的方法,使用R软件进行实际应用.
科学领域:
- 生物统计学 生物统计学
- 医疗信息学 医疗信息学
- 数据科学数据科学数据科学
背景情况:
- 在医学研究中,对纵向特征的聚类对于识别疾病发展轨迹至关重要.
- 整合多个纵向特征通过结合更多信息来增强聚类,可能揭示共存的模式和更深入的生物学见解.
- 对于实施和评估医疗数据集中的多个纵向特征的集群分析,存在有限的实际指导.
研究的目的:
- 概述用于聚类多个纵向特征的常用方法.
- 提供使用R软件实现的实用指导.
- 在医疗环境中比较不同集群方法的性能.
主要方法:
- 基于模型 (频率主义和贝叶斯主义) 和基于算法的方法的概述,用于聚类多个纵向特征.
- 强调使用R软件的应用和实施.
- 使用现实生活和模拟医疗数据集进行比较性绩效评估.
主要成果:
- 对多个纵向特征进行各种聚类方法的比较.
- 对不同数据集类型 (现实和模拟) 的性能评估.
- 为应用研究人员确定实践指南.
结论:
- 该研究为研究人员应用聚类方法对多个纵向特征提供了实际指导.
- 它为应用研究人员提供了建议,以及该领域未来的研究方向.
- 对比分析有助于选择适当的方法来分析医疗数据.
相关概念视频
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Comparing the Survival Analysis of Two or More Groups
218
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
218
Statistical Analysis: Overview
6.6K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.6K
Extraction: Advanced Methods
481
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
481
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K


