在高维时间序列数据中检测异常与缩放的布雷格曼分歧
Yunge Wang1, Lingling Zhang2, Tong Si3
1Department of Mathematics and Statistics, Saint Louis University, St. Louis, MO 63103, USA.
概括
这项研究引入了一种新的异常检测算法,通过解决无边界性问题,有效处理高维数据. 这种新方法改善了复杂数据集中的异常识别.
科学领域:
- 数据科学数据科学数据科学
- 机器学习 机器学习
- 统计 统计 统计 统计
背景情况:
- 异常检测识别出偏离正常行为的数据点,具有广泛的应用.
- 高维数据给现有的异常检测算法带来了挑战.
- 无约束最小正方形重要度匹配 (uLSIF) 方法在某些高维场景中与无限制性作斗争.
研究的目的:
- 为设计用于高维数据的新型异常检测算法提出建议.
- 为了克服uLSIF等现有方法所遇到的无边界性问题.
- 提高复杂数据集中异常检测的准确性和适用性.
主要方法:
- 开发了一个规模化的基于Bregman分歧的异常检测算法.
- 纳入了参数学习的最小绝对偏差和最小平方损失.
- 在合成和现实世界的高维时间序列数据集上评估了算法.
主要成果:
- 拟议的算法有效地解决了无界问题.
- 在检测高维时间序列数据中的异常方面表现出卓越的性能.
- 在比较分析中表现优于其他基于密度比率估计的异常检测方法.
结论:
- 基于Bregman分歧的缩放算法是用于高维数据中异常检测的强有力的解决方案.
- 这种方法比现有技术有了显著的改进,特别是在具有挑战性的数据环境中.
- 算法的对现实世界数据集的有效性验证了它的实际实用性.
更多相关视频
13:44Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
42.7K
07:59Author Spotlight: Alignment of Synchronized Time-Series Data Using the Characterizing Loss of Cell Cycle Synchrony Model for Cross-Experiment Comparisons
Published on: June 9, 2023
1.3K
相关概念视频
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K
What Are Outliers?
3.6K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
3.6K
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
Mean Absolute Deviation
2.6K
The mean absolute deviation is also a measure of the variability of data in a sample. It is the absolute value of the average difference between the data values and the mean.
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
2.6K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Random Error
799
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
799
