Outlier Detection in Urban Air Quality Sensor Networks
V M van Zoest1, A Stein1, G Hoek2
11Faculty of Geo-Information Science and Earth Observation (ITC), University of Twente, PO Box 217, 7500 AE Enschede, The Netherlands.
Water, Air, and Soil Pollution
|March 23, 2018
Summary
A new method effectively identifies outliers in low-cost urban air quality sensor data, crucial for accurate nitrogen dioxide (NO2) monitoring. This approach accounts for the unique spatio-temporal variations in city environments.
Area of Science:
- Environmental Science
- Atmospheric Chemistry
- Sensor Technology
Background:
- Low-cost urban air quality sensor networks are vital for assessing air pollutant spatio-temporal variability.
- These sensors often produce erroneous data, including outliers, which challenge traditional detection methods due to urban environmental dynamics.
Purpose of the Study:
- To develop and validate a novel outlier detection method for nitrogen dioxide (NO2) concentrations from low-cost urban sensors.
- To address the limitations of existing methods in handling the complex spatio-temporal variations of urban air quality data.
Main Methods:
- A spatio-temporal classification approach was developed, dividing a year of hourly NO2 data into 16 classes.
- Classes were defined by station type (urban background vs. traffic), day type (weekday vs. weekend), and time of day.
- Outliers were identified within each class using statistical analysis based on the truncated normal distribution of NO2 observations.
Main Results:
- The novel method identified 0.1-0.5% of outliers in the Eindhoven low-cost air quality sensor network data.
- Detected outliers may represent measurement errors or significant air pollution events, requiring expert evaluation for definitive classification.
Conclusions:
- The proposed spatio-temporal classification method effectively detects outliers in urban NO2 data from low-cost sensors.
- This method preserves the inherent spatio-temporal variability crucial for accurate air quality analysis in urban settings.
Related Concept Videos
What Are Outliers?
5.2K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
5.2K
Outliers and Influential Points
6.3K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
6.3K
Quantifying and Rejecting Outliers: The Grubbs Test
4.2K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
4.2K
Protein Networks
4.6K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.6K
Network Covalent Solids
16.2K
Network covalent solids contain a three-dimensional network of covalently bonded atoms as found in the crystal structures of nonmetals like diamond, graphite, silicon, and some covalent compounds, such as silicon dioxide (sand) and silicon carbide (carborundum, the abrasive on sandpaper). Many minerals have networks of covalent bonds.
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...
16.2K
Quality Control
2.9K
Quality control is one of the three cyclical quality assurance activities that help keep a system under statistical control. Typical quality control activities include creating quality control charts, conducting proficiency testing, and documenting and archiving results.
Quality control helps track data, visualize trends, and identify variations, making it easier to detect deviations that may affect the accuracy of an analysis. One way to do this is by generating a quality control chart, which...
Quality control helps track data, visualize trends, and identify variations, making it easier to detect deviations that may affect the accuracy of an analysis. One way to do this is by generating a quality control chart, which...
2.9K


