一种检测线性循环非参数回归中的异常值的方法
Sümeyra Sert1, Filiz Kardiyen2
1Department of Statistics, Selcuk University, Selcuklu, Konya, Turkey.
PloS one
|June 12, 2023
概括
本研究介绍了一种强大的异常值检测方法,用于具有异常值的线性循环回归. 循环中位数方法有效处理受污染的数据,特别是在更大的样本大小和更高的一致性的情况下.
科学领域:
- 统计 统计 统计 统计
- 数据科学数据科学数据科学
背景情况:
- 非参数回归对于建模复杂关系至关重要.
- 响应变量的异常值可以显著扭曲回归结果.
- 线性循环回归用于线性和循环组成部分的数据.
研究的目的:
- 为非参数线性循环回归提出一个强大的异常值检测方法.
- 为了应对响应变量中异常值所带来的挑战,当残留量遵循包装-考奇分布时.
- 评估在各种条件下提出的方法的性能.
主要方法:
- 开发一种强大的异常值检测方法,利用圆形中位数.
- 纳达拉亚-沃森和局部线性回归的应用对于非参数的匹配.
- 通过真实数据集分析和全面的模拟研究进行绩效评估.
主要成果:
- 拟议的方法显示出强大的性能,特别是在中高污染水平下.
- 随着样本大小和数据均性的增加,方法的有效性会提高.
- 当局线性估计在响应变量中存在异常值时,其表现优于Nadaraya-Watson.
结论:
- 基于循环中位数的异常值检测方法为带有污染数据的非参数线性循环回归提供了可靠的解决方案.
- 选择回归拟合方法 (局部线性估计与Nadaraya-Watson) 在异常值存在时会影响性能.
- 该研究强调了样本大小和数据同质性对于稳健的统计建模的重要性.
相关概念视频
Outliers and Influential Points
4.1K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.1K
Quantifying and Rejecting Outliers: The Grubbs Test
1.7K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.7K
What Are Outliers?
3.9K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
3.9K
Detection of Gross Error: The Q Test
6.3K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.3K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Modified Boxplots
9.8K
A standard box and whisker plot informs us about the spread of the data in a given sample. One can identify the minimum value, maximum value, first quartile value, second quartile or median value, and third quartile.
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
9.8K


