H-NGPCA:数据流的层次分类,具有可适应的集群数和可适应的维度
Nico Migenda1, Ralf Möller2, Wolfram Schenck1
1Center for Applied Data Science Gütersloh, Bielefeld University of Applied Sciences and Arts, Bielefeld, Germany.
我们介绍了H-NGPCA,这是一种用于数据流的新型层次聚类算法. 这种自适应算法通过整合在线维度控制和单位增长来实现与离线方法相比较的高精度.
科学领域:
- 机器学习 机器学习
- 数据挖掘 数据挖掘
- 人工智能的人工智能
背景情况:
- 传统的集群算法与动态数据流作斗争.
- 现有的在线方法往往缺乏单位数和维度的适应性.
研究的目的:
- 开发一个对数据流的层次聚类算法,具有自适应单位数量增长和局部维度控制.
- 结合基于中心体,基于模型和层次的集群特征.
主要方法:
- H-NGPCA建立了当地主要组件分析 (PCA) 单元的层次结构.
- 单位是通过基于神经网络的在线PCA更新的超圆形.
- 神经气处理单元重新定位,一个分割标准管理单元创建.
- 每个单位自行确定其维度.
主要成果:
- H-NGPCA以适应性单元数优于竞争中的在线算法.
- 通过最先进的离线方法实现了竞争性表现.
- 证明了高准确性,平均规范化相互信息 (NMI) 为0.87和集群指数 (CI) 为0.26.
结论:
- H-NGPCA为数据流提供了卓越的在线适应性.
- 该算法在动态环境中实现了离线级准确性.
更多相关视频
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
06:01Visualization and Quantification of High-Dimensional Cytometry Data using Cytofast and the Upstream Clustering Methods FlowSOM and Cytosplore
Published on: December 12, 2019
相关概念视频
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Collisions in Multiple Dimensions: Introduction
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Rapidly Varying Flow
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Dimensional Analysis
In fluid mechanics, dimensional...
