Time series AQI forecasting using Kalman-integrated Bi-GRU and Chi-square divergence optimization
Narmeen Fatima1,2, Samia Nawaz Yousafzai3, Nadhem Nemri4
1Applied INTelligence Lab (AINTLab), Seoul, 05006, Republic of Korea.
Scientific Reports
|August 9, 2025
Summary
This study introduces a novel deep learning framework for accurate air quality index (AQI) forecasting, improving public health by addressing data uncertainty and missing values in environmental monitoring.
Area of Science:
- Environmental Science
- Data Science
- Machine Learning
Background:
- Air pollution poses a significant global health risk, necessitating reliable air quality index (AQI) prediction systems.
- Current AQI forecasting models struggle with missing data, high variability, and distributional uncertainty, limiting their effectiveness.
- Accurate AQI forecasting is crucial for public health protection and environmental policy development.
Purpose of the Study:
- To develop a novel deep learning framework for robust AQI time-series forecasting.
- To address limitations in existing models, including missing data, data variability, and distributional uncertainty.
- To improve the accuracy and reliability of AQI predictions for enhanced environmental monitoring.
Main Methods:
- Integration of Kalman Attention with a Bi-Directional Gated Recurrent Unit (Bi-GRU) for dynamic uncertainty handling and temporal feature weighting.
- Incorporation of a Chi-square Divergence-based regularization term to minimize distributional mismatch between predicted and actual pollutant levels.
- Imputation of missing values using pollutant-specific ARIMA models to preserve time-dependent trends.
Main Results:
- The proposed framework demonstrated significant improvements over baseline models (LSTM, CNN-LSTM) in AQI forecasting.
- Achieved a high R-squared value of 0.96794, indicating strong model performance.
- Reported a Mean Squared Error (MSE) of 4.11×10⁻⁵ and Mean Absolute Error (MAE) of 0.000423, signifying high prediction accuracy.
Conclusions:
- The novel deep learning framework effectively addresses key challenges in AQI forecasting, including uncertainty, distributional alignment, and missing data.
- The integrated architecture provides a scalable solution for environmental monitoring and supports evidence-based policy decisions.
- This research advances the field of AQI prediction, offering a more robust and reliable tool for safeguarding public health from air pollution.
Related Concept Videos
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
712
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
712
Chi-square Distribution
4.7K
How does one determine if bingo numbers are evenly distributed or if some numbers occurred with a greater frequency? Or if the types of movies people preferred were different across different age groups or if a coffee machine dispensed approximately the same amount of coffee each time. These questions can be addressed by conducting a hypothesis test. One distribution that can be used to find answers to such questions is known as the chi-square distribution. The chi-square distribution has...
4.7K
Calibration Curves: Linear Least Squares
2.2K
A calibration curve is a plot of the instrument's response against a series of known concentrations of a substance. This curve is used to set the instrument response levels, using the substance and its concentrations as standards. Alternatively, or additionally, an equation is fitted to the calibration curve plot and subsequently used to calculate the unknown concentrations of other samples reliably.
For data that follow a straight line, the standard method for fitting is the linear...
For data that follow a straight line, the standard method for fitting is the linear...
2.2K
Chi-square Analysis
38.7K
The chi-square test is a statistical hypothesis test. It is used to check whether there is a significant difference between an expected value and an observed value. In the context of genetics, it enables us to either accept or reject a hypothesis, based on how much the observed values deviate from the expected values.
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
38.7K
Maxwell-Boltzmann Distribution: Problem Solving
1.7K
Individual molecules in a gas move in random directions, but a gas containing numerous molecules has a predictable distribution of molecular speeds, which is known as the Maxwell-Boltzmann distribution, f(v).
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
1.7K
Quantitative Analysis
606
Quantitative analysis is a technique for measuring the amount of specific constituents in a sample. When the sample's composition is unknown, qualitative analysis is performed first to identify its components, which ensures that the correct substances are measured during the quantitative phase.
In quantitative analysis, two key measurements are made: the sample quantity and a property proportional to the amount of the analyte (the substance being analyzed). This forms the basis of the...
In quantitative analysis, two key measurements are made: the sample quantity and a property proportional to the amount of the analyte (the substance being analyzed). This forms the basis of the...
606


