聚类对计数数据的统计意义
Yifan Dai1, Di Wu1,2, Yufeng Liu3
1Department of Biostatistics, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, USA.
Biometrics
|September 19, 2025
概括
我们介绍了SigClust-DEV,这是一种用于评估计数数据中的集群意义的新方法,其性能优于现有的方法. 该工具通过解决统计不确定性来增强基因组学和医疗保健数据中的子组识别.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 统计遗传学 统计遗传学
背景情况:
- 聚类对于确定生物医学研究中的子组至关重要,但现有的方法往往忽视统计不确定性,导致虚假的聚类.
- 聚类的统计学意义 (SigClust) 方法在高维数据中评估聚类意义,但仅限于连续数据,并且可能缺乏非高斯分布的权力.
研究的目的:
- 开发一种新的方法,SigClust-DEV,用于评估集群的统计意义,特别是在离散计数数据中.
- 解决现有的SigClust方法在处理计数数据和非高斯分布方面的局限性.
主要方法:
- SigClust-DEV是为了评估高维数计数数据中的集群意义而开发的.
- 进行了广泛的模拟,将SigClust-DEV与各种计数分布的现有SigClust变异进行比较.
- 该方法应用于单细胞RNA测序 (scRNA) 数据和电子健康记录 (EHR) 进行现实世界验证.
主要成果:
- 与现有的SigClust方法相比,SigClust-DEV在各种计数分布的模拟中表现出更高的性能.
- 该方法在Hydra scRNA数据中成功识别了有意义的潜伏细胞类型.
- 在癌症EHR数据中,SigClust-DEV有效地识别了显著的患者子组.
结论:
- SigClust-DEV是一种强大且在统计学上可靠的方法,用于评估计数数据中的集群意义.
- 该方法增强了复杂的生物医学数据集 (如scRNA和EHR数据) 中的子组发现.
- SigClust-DEV克服了以前方法的局限性,为离散数据分析提供了更好的统计能力和准确性.
更多相关视频
13:55Combined Immunofluorescence and DNA FISH on 3D-preserved Interphase Nuclei to Study Changes in 3D Nuclear Organization
Published on: February 3, 2013
18.9K
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
493
相关概念视频
Statistical Significance
21.1K
Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
21.1K
Cluster Sampling Method
14.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
14.0K
Introduction to Test of Independence
2.9K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.9K
Test for Homogeneity
2.4K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.4K
Significance Testing: Overview
11.5K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
11.5K
Determination of Expected Frequency
2.5K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.5K
