Related Experiment Video
Updated: Jun 11, 2025

Automated, High-resolution Mobile Collection System for the Nitrogen Isotopic Analysis of NOx
Published on: December 20, 2016
A probabilistic framework for identifying anomalies in urban air quality data.
Priti Khatri1,2, Kaushlesh Singh Shakya1,2, Prashant Kumar3,4
1Academy of Scientific & Innovative Research (AcSIR), Ghaziabad, 201002, India.
This study introduces a new method to detect and remove errors in air quality data, focusing on particulate matter (PM2.5 and PM10) in Delhi. Improving data quality is crucial for accurate health and environmental protection decisions.
Area of Science:
- Environmental Science
- Data Science
- Atmospheric Chemistry
Background:
- Air quality data requires systematic processing to be useful for decision-making.
- Ground-based monitoring data often contains outliers, which can be error-based or event-based.
- Error-based outliers (e.g., instrument malfunction, sensor drift) are noise and must be removed, unlike meaningful event-based outliers.
Purpose of the Study:
- To develop and validate a robust methodology for detecting error-based outliers in air quality data.
- To specifically target particulate matter (PM2.5 and PM10) data from monitoring sites in Delhi.
- To improve the reliability of air quality data for subsequent analysis and modeling.
Main Methods:
- Developed a non-linear filtering approach to model air quality data.
- Calculated residuals between observed and predicted values.
- Utilized Z-scores to assess residual probabilities and flag outliers based on a sensitivity-tested threshold.
Main Results:
- Identified four distinct types of error-based outliers in PM2.5 and PM10 data: extreme values, constant readings/low variance, periodic self-calibration issues, and PM2.5/PM10 ratio anomalies.
- Successfully flagged outliers using the developed Z-score based probability method.
- Demonstrated the method's effectiveness on data with less than 5% missing values.
Conclusions:
- The proposed methodology effectively detects error-based outliers in air quality data.
- Refined data quality is essential for accurate air quality modeling and reliable statistical/machine learning applications.
- This approach enhances the integrity of environmental data for informed decision-making.
More Related Videos
Related Concept Videos
Random Error
Steps in Outbreak Investigation
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Quantifying and Rejecting Outliers: The Grubbs Test
Detection of Gross Error: The Q Test
Mechanistic Models: Compartment Models in Individual and Population Analysis

