离散和平衡的光谱集群与可扩展性
IEEE transactions on pattern analysis and machine intelligence
|September 5, 2023
概括
本研究介绍了离散和平衡的光谱集群与可扩展性 (DBSC),这是一种改进数据集群的新方法. DBSC解决了现有的光谱聚类技术的局限性,为大型数据集提供了更好的性能和可扩展性.
科学领域:
- 机器学习 机器学习
- 数据挖掘 数据挖掘
- 计算机视觉 计算机视觉
背景情况:
- 频谱聚类 (SC) 是广泛使用的,但面临的挑战是从两阶段方法中获得次优解决方案.
- 现有的SC方法难以保持集群平衡,并且对于大型数据集而言,计算成本昂贵.
研究的目的:
- 开发一种新的离散和平衡的光谱集群与可扩展性 (DBSC) 模型.
- 克服现有的SC方法的局限性,包括低于最佳的解决方案,缺乏平衡属性和可扩展性差.
主要方法:
- 将连续放松和离散集群指标矩阵的集成学习整合到单个步骤中.
- 纳入了基于的策略,以提高对大规模数据集的可扩展性.
- 通过保持大约相同的集群大小,实现了软平衡的集群.
主要成果:
- 与现有方法相比,DBSC模型表现出优越的集群和平衡性能.
- 在CMUPIE数据上的聚类准确性比最先进的方法显著提高了17.93%.
- 成功解决了次优解决方案,集群不平衡和计算成本的问题.
结论:
- DBSC为光谱聚类提供了更有效和高效的方法.
- 该模型的集成,平衡和可扩展的设计使其适合于现实世界的大规模应用.
- DBSC代表了光谱聚类技术的重大进步.
更多相关视频
07:11ARL Spectral Fitting as an Application to Augment Spectral Data via Franck-Condon Lineshape Analysis and Color Analysis
Published on: August 19, 2021
2.5K
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
7.0K
相关概念视频
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Extraction: Partition and Distribution Coefficients
2.5K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
2.5K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Discrete Fourier Transform
320
The Discrete Fourier Transform (DFT) is a fundamental tool in signal processing, extending the discrete-time Fourier transform by evaluating discrete signals at uniformly spaced frequency intervals. This transformation converts a finite sequence of time-domain samples into frequency components, each representing complex sinusoids ordered by frequency. The DFT translates these sequences into the frequency domain, effectively indicating the magnitude and phase of each frequency component present...
320
IR Spectrum Peak Splitting: Symmetric vs Asymmetric Vibrations
1.1K
Identical bonds within a polyatomic group can stretch symmetrically (in-phase) or asymmetrically (out-of-phase). Similar to hydrogen bonding, these vibrations also influence the shape of the IR peak. Generally, asymmetric stretching frequencies are higher than symmetric stretching frequencies. For example, primary amines exhibit two distinct IR peaks between 3300–3500 cm−1 corresponding to the symmetric and asymmetric N-H stretching, while secondary amines exhibit a single...
1.1K
Construction of Frequency Distribution
7.8K
A frequency distribution table can be constructed using the steps given below.
First, make a table with two columns—one with the title of the data that needs to be organized, and the other column for frequency. [Draw a third column for tally marks if needed]. Then, take a look at the items given in the data set and decide if an ungrouped frequency distribution table or a grouped frequency distribution table would be more suitable. If there are large sets of different values, then it is...
First, make a table with two columns—one with the title of the data that needs to be organized, and the other column for frequency. [Draw a third column for tally marks if needed]. Then, take a look at the items given in the data set and decide if an ungrouped frequency distribution table or a grouped frequency distribution table would be more suitable. If there are large sets of different values, then it is...
7.8K
