基础真理聚类不是最优的聚类
Lucia Absalom Bautista1, Timotej Hrga2, Janez Povh3,4
1University of Sevilla, C. San Fernando 4, Seville, 41004, Spain.
Scientific reports
|March 18, 2025
概括
最佳的数据聚类解决方案往往不同于基本真相,但可以产生更高的内在质量. 当地面真理集群是分开得很好的凸起形状时,对齐性会得到改善.
科学领域:
- 数据科学数据科学数据科学
- 计算统计学 计算统计学
- 机器学习 机器学习
背景情况:
- 数据聚类至关重要,但具有挑战性.
- 最小平方和集群 (MSSC) 旨在将点到中心距离最小化.
- MSSC是NP-hard,但对于最佳解决方案存在解决者.
研究的目的:
- 使用SOS-SDP解决器获得和评估最佳的MSSC解决方案.
- 在各种数据集上比较最佳集群与基准真相集群.
- 通过外部和内在的措施来评估最佳集群的质量.
主要方法:
- 利用了SOS-SDP解决器,这是一个基于半确定的编程的分支和结合算法.
- 获得了最佳的MSSC解决方案,用于已知基础真相的各种数据集.
- 评估的集群对齐与六个外部和三个内在的措施.
主要成果:
- 最佳的集群常常与基础真相的集群有所不同.
- 最佳的集群通常表现出与基本真相相比更高的内在质量.
- 当地面真相集群是与圆形相似的凸形状分离得很好时,观察到的高度对齐.
结论:
- 最佳的MSSC解决方案可能并不总是与人类定义的基本事实相匹配.
- 在数学上最佳的解决方案中,内在的集群质量可能更高.
- 数据集群的几何性质影响了最佳和基本真相解决方案之间的协议.
相关概念视频
Quantifying and Rejecting Outliers: The Grubbs Test
1.4K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.4K
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K
Sampling Plans
162
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
162
Detection of Gross Error: The Q Test
5.0K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
5.0K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Comparing the Survival Analysis of Two or More Groups
110
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
110


