发现异常值对时间数据集中的集群演变的影响:经验分析
Muhammad Atif1, Muhammad Farooq1, Muhammad Shafiq2
1Department of Statistics, University of Peshawar, Peshawar, Pakistan.
Scientific reports
|December 27, 2024
概括
异常值在时间数据中显著影响集群演变. 适当处理异常值对于可靠的流集群结果和开发强大的算法至关重要.
科学领域:
- 数据科学数据科学数据科学
- 机器学习 机器学习
- 人工智能的人工智能
背景情况:
- 监测时间数据集中的集群转换显示了数据动态.
- 了解集群演变对于知情数据分析和决策至关重要.
- 集群演变包括由于新数据点而随着时间的推移而发生的集群变化.
研究的目的:
- 调查异常值对时间数据集集群演变的影响.
- 分析异常值如何影响集群过渡和稳定性随时间推移.
- 评估异常值处理的必要性,以实现准确的流集群.
主要方法:
- 使用生存率和历史成本函数来量化异常效应.
- 在相继的时间点之间跟踪集群之间的数据点移动.
- 分类集群解决方案的变化变成了外部和内部的过渡.
主要成果:
- 发现异常值对集群演变有重大影响.
- 该研究确定了异常值对集群变化的特定影响.
- 在未处理的异常值存在时,观察到不一致的集群结果.
结论:
- 适当的异常值处理技术对于可靠的流集群至关重要.
- 结果为提高流集群算法的稳定性提供了见解.
- 这项研究指导着开发更准确,更可靠的时间数据分析方法.
相关概念视频
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
What Are Outliers?
3.6K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
3.6K
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K
Detection of Gross Error: The Q Test
5.6K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
5.6K
Modified Boxplots
9.1K
A standard box and whisker plot informs us about the spread of the data in a given sample. One can identify the minimum value, maximum value, first quartile value, second quartile or median value, and third quartile.
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
9.1K
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K


