在基于模型的聚类中处理斜率和定向尾巴
Cristina Tortora1, Antonio Punzo2, Brian C Franczak3
1Department of Mathematics and Statistics, San José State University, One Washington square, San José, California 95192 USA.
概括
本研究引入了新的基于模型的集群方法,使用转换的多个尺度受污染的正常分布 (MSCN). 这些方法通过处理偏差数据和更有效地检测异常值来改善数据分析.
科学领域:
- 统计 统计 统计 统计
- 数据挖掘 数据挖掘
- 机器学习 机器学习
背景情况:
- 基于模型的聚类对于识别数据模式至关重要.
- 标准方法与偏斜或重尾数据集群作斗争.
- 在复杂的数据集中,异常值的检测仍然是一个挑战.
研究的目的:
- 引入使用转换MSCN分布的新型集群模型.
- 增强集群形状的灵活性 (斜度,曲度).
- 启用组件智能和定向异常值检测.
主要方法:
- 开发了基于观察到数据的组件智能转换的两个模型.
- 多个尺度受污染的正常分布 (MSCN) 的使用混合物.
- 采用主要组件的内置定向异常值检测.
主要成果:
- 拟议的基于MSCN的集群提供灵活的集群形状.
- 实现了有效的组件智能和定向异常值检测.
- 与全球/组件智能异常值检测方法相比,已证明的优势.
结论:
- 基于MSCN的方法为复杂的数据提供了强大的集群.
- 在各种尺寸中提供卓越的异常值检测能力.
- 在特定的实践集群场景中优于现有方法.
更多相关视频
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
7.1K
08:38A System for Tracking the Dynamics of Social Preference Behavior in Small Rodents
Published on: November 21, 2019
7.7K
相关概念视频
Types of Skewness
12.6K
If the frequency distribution of a data set is more inclined towards smaller or larger values, the distribution is said to be skewed. If data values are skewed to the right, then the distribution is called positively skewed. Conversely, if the plot is skewed to the left, the distribution is called negatively skewed.
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
12.6K
Skewness
12.7K
The measures of central tendency calculated from a data set may not reveal much about its intrinsic distribution. If a plot is made of the data set’s values, the mean and the median may not only differ, but also the plot may have more values on one side of the central tendencies. Such a data set is said to be skewed towards that side.
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
12.7K
Microsoft Excel: Finding Central Tendency, Skew, and Kurtosis
312
Central tendency refers to the central point or typical value of a dataset. It summarizes the data set with a single value that represents the center of its distribution. The three main measures of central tendency are:
Mean: The arithmetic average of all data points. It is calculated by adding all the values together and dividing by the number of values. The mean is sensitive to extreme values (outliers).
Median: The middle value when the data points are arranged in ascending or descending...
Mean: The arithmetic average of all data points. It is calculated by adding all the values together and dividing by the number of values. The mean is sensitive to extreme values (outliers).
Median: The middle value when the data points are arranged in ascending or descending...
312
Distributions to Estimate Population Parameter
4.3K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.3K
Cluster Sampling Method
12.8K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.8K
Modified Boxplots
10.1K
A standard box and whisker plot informs us about the spread of the data in a given sample. One can identify the minimum value, maximum value, first quartile value, second quartile or median value, and third quartile.
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
10.1K
