用多维缩放进行集群的统计意义
Hui Shen1, Shankar Bhamidi1, Yufeng Liu2
1Department of Statistics and Operations Research, University of North Carolina at Chapel Hill, U.S.A.
概括
一种新方法增强了对高维数据的集群意义测试. 这种方法使用多维缩放 (MDS) 与不相似度矩阵,在原始数据不可用时提高可靠性.
科学领域:
- 数据科学数据科学数据科学
- 计算统计学 计算统计学
- 生物信息学是一种生物信息学.
背景情况:
- 集群对于探索性数据分析至关重要,但评估集群可靠性是具有挑战性的.
- 集群的现有统计学意义 (SigClust) 方法可能会在某些高维,低样本大小的场景中失败,或者只有不相似度矩阵可用.
研究的目的:
- 开发一种新的SigClust方法,适用于当研究人员只能访问不相似度矩阵时.
- 在特定的高维数据场景中解决原始SigClust方法的局限性.
主要方法:
- 提出了一种新的SigClust方法,利用多维缩放 (MDS).
- MDS用于从不相似矩阵创建低维表示.
- 然后,SigClust应用于这些低维的MDS产生的空间.
主要成果:
- 基于MDS的SigClust有效评估集群的统计意义.
- 这种方法绕过了高维空间中的参数估计挑战.
- 它保留了MDS生成的低维空间中的基本集群结构.
结论:
- 拟议的基于MDS的SigClust是评估集群意义的强大和适用的工具.
- 它将SigClust的实用性扩展到有限的数据访问情况 (仅不相似度矩阵).
- 通过模拟和现实世界的数据应用来证明有效性.
相关概念视频
Statistical Significance
20.1K
Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
20.1K
One-Way ANOVA: Equal Sample Sizes
3.2K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.2K
One-Way ANOVA: Unequal Sample Sizes
5.7K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.7K
Spearman's Rank Correlation Test
685
Spearman's rank correlation test, also known as Spearman's rho, is a nonparametric method for assessing the strength and direction of association between two variables. This test is particularly valuable when the data distribution is unknown or when the assumption of normality does not hold. Named after the English psychologist and statistician Dr. Charles Edward Spearman, it serves as the nonparametric counterpart to Pearson's correlation coefficient.
Spearman's test calculates...
Spearman's test calculates...
685
Significance Testing: Overview
3.3K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
3.3K
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K


