流动集群数据的实时回归分析,其中可能存在异常数据批次
Lan Luo1, Ling Zhou2, Peter X-K Song3
1Department of Statistics and Actuarial Science, University of Iowa.
Journal of the American Statistical Association
|September 29, 2023
概括
本研究引入了一种新的可再生二次推理函数 (RenewQIF),用于分析流数据. 这种方法有效地更新统计模型而不需要历史原始数据,证明对相关结果有效.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 分析具有相关结果的流数据集,如纵向和集群数据,会带来重大的计算和统计挑战.
- 现有的方法通常需要重新处理所有历史数据以进行更新,从而导致效率低下.
- 对于动态数据集的高效增量学习算法的需求至关重要.
研究的目的:
- 开发一种增量学习算法,用于分析具有相关结果的流数据集.
- 提出一种新的可再生二次推理函数 (RenewQIF) 方法,用于高效的参数估计.
- 引入一个顺序的合适性测试,用于诊断流数据中的回归系数均性.
主要方法:
- 开发了一个可再生的二次推理函数 (RenewQIF) 用于增量学习.
- 采用可再生估计和增量推断范式,更新参数与当前数据和历史总结统计数据.
- 将RenewQIF与离线二次推理函数 (QIF) 和概括估计方程 (GEE) 方法进行比较.
- 提出了一种顺序的适合性测试,用于同质性假设诊断.
- 使用扩展的 Spark 的 Lambda 架构实现了该方法.
主要成果:
- 从理论和数值上证明,与离线方法相比,可再生能源程序提供了统计和计算效率.
- 拟议的顺序性合适性测试有效地选异常数据批,并诊断出违反同质性假设的情况.
- 该方法通过广泛的模拟研究和对汽车事故数据的现实世界分析,成功地展示了该方法.
结论:
- 拟议的RenewQIF方法提供了一种高效和统计学上合理的方法来分析流数据,并提供相关的结果.
- 综合数据质量诊断工具在动态环境中提高了统计推理的可靠性.
- 这项工作为大数据应用中的实时统计分析提供了可扩展的解决方案.
相关概念视频
Steps in Outbreak Investigation
152
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
152
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Interpreting Run Charts
126
Run charts, essentially line graphs plotted over time, serve as fundamental yet effective tools for process analysis. They chronicle data sequentially, facilitating the identification of trends, shifts, or cyclical movements. This graphical representation is instrumental in determining whether a process is stable or exhibits signs of potential instability indicative of special cause variation. In the healthcare domain, run charts depict infection rates over time, enabling hospitals to monitor...
126
Random Error
922
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
922
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K


