分区的低最佳比较
Jonathon J O'Brien1, Michael T Lawson2, Devin K Schweppe1
1Department of Cell Biology, Harvard Medical School, Boston, MA 02115, USA.
概括
最佳的聚类和分类解决方案分歧,即使是已知的数据模型. 标准验证措施不能保证最佳的集群性能,因此需要采用替代性评估方法进行后期解释.
科学领域:
- 计算统计的计算统计.
- 机器学习 机器学习
- 数据挖掘是一种数据挖掘.
背景情况:
- 分类和聚类是不同的数据分析任务,通常通过先前标签的可用性来区分.
- 理论分析表明,集群的最佳解决方案并不总是与最佳分类保持一致,特别是当数据生成模型已知时.
研究的目的:
- 探索最佳集群和分类之间的分歧.
- 确定标准内部和外部验证措施在保证最佳集群性能方面的局限性.
- 为集群推低于最佳的评估策略,提供有价值的后期解释.
主要方法:
- 理论探讨聚类和分类最佳性之间的关系.
- 分析标准的内部和外部验证指数.
- 开发和推用于集群绩效的替代评估指标.
主要成果:
- 没有任何标准的内部或外部验证措施可以确保与最佳集群对应.
- 基于对联的指数为聚类提供了明确的概率解释.
- 基于三元组的指数揭示了更高层次的数据结构,而来自等级聚类树状图的ROC曲线提供了细微的见解.
结论:
- 标准验证指标不足以保证最佳集群结果.
- 低于最佳的评估方法,包括对和三重指数,对于后期集群解释是有价值的.
- 图形方法,如来自树图的ROC曲线,为理解集群结果提供了比单个数字摘要更丰富的信息.
相关概念视频
Extraction: Partition and Distribution Coefficients
1.8K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
1.8K
Multiple Comparison Tests
3.8K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.8K
Comparing the Survival Analysis of Two or More Groups
146
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
146
Compacting Factor test
115
The compacting factor test is a method used to assess the workability of concrete. It is especially suitable for concrete mixes containing aggregates up to one and a half inches in size. This test involves specialized equipment consisting of two truncated cone-shaped hoppers and a cylinder, all with polished interior surfaces to minimize friction.
The procedure begins by placing concrete into the upper hopper without any compaction. Once filled, the bottom door of this hopper is opened,...
The procedure begins by placing concrete into the upper hopper without any compaction. Once filled, the bottom door of this hopper is opened,...
115
Friedman Two-way Analysis of Variance by Ranks
144
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
144
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K


