一个概率框架,用于识别城市空气质量数据中的异常
Priti Khatri1,2, Kaushlesh Singh Shakya1,2, Prashant Kumar3,4
1Academy of Scientific & Innovative Research (AcSIR), Ghaziabad, 201002, India.
Environmental science and pollution research international
|October 2, 2024
概括
这项研究引入了一种新方法来检测和删除空气质量数据中的错误,重点关注德里的颗粒物 (PM2.5和PM10). 提高数据质量对于准确的健康和环境保护决策至关重要.
科学领域:
- 环境科学 环境科学
- 数据科学数据科学数据科学
- 大气化学 大气化学
背景情况:
- 空气质量数据需要系统处理才能对决策有用.
- 基于地面的监测数据通常包含异常值,这些异常值可以基于错误或基于事件.
- 基于错误的异常值 (例如仪器故障,传感器漂移) 是噪声,必须被删除,与基于有意义事件的异常值不同.
研究的目的:
- 开发和验证一个可靠的方法来检测空气质量数据中的基于错误的异常值.
- 专门针对来自德里监测站点的颗粒物 (PM2.5和PM10) 数据.
- 提高空气质量数据的可靠性,用于随后的分析和建模.
主要方法:
- 开发了一种非线性过方法来建模空气质量数据.
- 观察值和预测值之间的计算余量.
- 使用Z-score来评估基于灵敏度测试值的剩余概率和标记异常值.
主要成果:
- 在PM2.5和PM10数据中确定了四种不同类型的基于错误的异常值:极端值,恒定读数/低方差,定期自我校准问题和PM2.5/PM10比率异常.
- 使用开发的基于Z分数的概率方法,成功标记了异常值.
- 证明了该方法在不到5%缺失值的数据上的有效性.
结论:
- 拟议的方法有效地检测空气质量数据中的基于错误的异常值.
- 精细的数据质量对于准确的空气质量建模和可靠的统计/机器学习应用至关重要.
- 这种方法增强了环境数据的完整性,以提供知情决策.
更多相关视频
相关概念视频
Random Error
843
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
843
Steps in Outbreak Investigation
108
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
108
Sampling Plans
169
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
169
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K
Detection of Gross Error: The Q Test
5.7K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
5.7K
Mechanistic Models: Compartment Models in Individual and Population Analysis
32
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
32


