0% found this document useful (0 votes)
31 views64 pages

Air Quality Impact and Machine Learning

The document discusses the critical issue of global climate change and its impact on human health and ecosystems, emphasizing the need for effective monitoring and prediction of air quality. It outlines various sources of air pollution, including industrial emissions and transportation, and highlights the significant health risks associated with poor air quality, leading to millions of premature deaths annually. Additionally, it explores the role of machine learning in improving air quality prediction and monitoring through advanced data analysis and real-time applications.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
31 views64 pages

Air Quality Impact and Machine Learning

The document discusses the critical issue of global climate change and its impact on human health and ecosystems, emphasizing the need for effective monitoring and prediction of air quality. It outlines various sources of air pollution, including industrial emissions and transportation, and highlights the significant health risks associated with poor air quality, leading to millions of premature deaths annually. Additionally, it explores the role of machine learning in improving air quality prediction and monitoring through advanced data analysis and real-time applications.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CHAPTER 1

INTRODUCTION

1.1 Introduction

Global climate is the major issue which is rising in 21st century affecting human
health and ecosystems. Monitoring and predicting of climate are very crucial to
identify its effects and to make a sustainable environment. Due to industrial growth
and urbanization, it leads to the environmental challenges. It has a very bad
influence on environment, physical and mental health and economic
development[1]. Industrial Pollution results to various respiratory and
cardiovascular diseases moreover it leads to reproductive diseases[2]. Air pollution
cause premature deaths by directly affecting the respiratory systems in children. It
increases the chances of asthma, heart issues, skin infections, eye diseases, throat
infections, lung cancer, etc. It can affect the children for impaired lung function and
cognitive developments [3].

1.2 Sources of pollution

The term “Air quality” means the health of the air or the condition of the air in the
environment. Concentration of pollutants and presence of harmful substance pollute
the air or decrease the air health. Pollutants like CO, CO2, NO2, NO, SO2, NH3,
PM2.5, PM10, O3 etc. influenced the air quality. These are originated from the
wildfires, volcanic eruptions, industrial emissions, vehicles, burning fossil fuels and
agriculture waste. It harms human health as well as the environment such as
harming rivers and vegetation [2][4]. Due to these dangerous pollutant levels and
poor-quality people are affected in different ways. Air quality is usually referred as
Air Quality Index (AQI). Almost 50% of the pollution in India comes from
industrial emissions, 27% of pollution comes from the vehicle emissions. rest of the
pollutions comes from the crop burning, Domestic cooking etc.[3]. All the
pollutants are integrated from the natural and the human-made sources that is
anthropogenic.

1
1.2.1 Anthropogenic Sources

a) Transportation: Public and private vehicle emissions powered by petrol ,


diesel and CNG emit the pollutants like CO, NOx and PM

b) Industrial Activities: Activities such as cement production, oil refining,


chemical manufacturing, bricks manufacturing etc. are the major
contributors which emit the gases like SO2, NOx etc.

c) Heating and Cooking: Burning coal, wood, biomass releases the


pollutants like PM, CO etc.

1.3 Influence on human and economic health

Due to the human health, it decreases the work effort which directly effects the
economy. So, government needs to focus on this issue to sustain human health,
economic health and ecosystems[3].

Table 1.1 Air quality index values and their levels

Air Quality Index Value Levels


0 - 50 Good
51 – 100 Moderate
101-150 Unhealthy for sensitive groups
151-200 Unhealthy
201- 300 Very Unhealthy
301 and above Hazardous

Air quality Index is divided into groups. First group of AQI index is between 0 to 50
which is good and tells that very low pollution. Second group is between 51 to 100
which is at moderate level. Third group is between 101- 150 which is unhealthy for
sensitive groups that means human with respiratory disease will have difficulty to
breathe. Fourth is 151-200 which is unhealthy for all beings. Fifth group is between 201-
300 which again is very unhealthy and cause diseases. AQI above 301 lies under sixth
group which is hazardous ad can cause serious health issues. More the AQI level more
polluted the air is [5].

According to WHO 99% of the population of the worlds was not taking right measures
in [Link] pollution cause 4.2 million premature deaths on the same year. From 4.2

2
million 68% of the premature deaths happened due to heart disease and stroke, 14%
were due to chronic obstructive pulmonary disease ,14% were due to acute lower
respiratory infections and 4% were due to lung cancers.

Premature deaths Reasons

4%
14%

14%

68%

Heart Disease & Strokes Obstructive pulmonary disease Respiratory Infections Lung Cancers

Figure 1.1 Premature Deaths and their reasons

1.4 Role of Machine learning in air quality

As air quality monitoring is very crucial to identify the pollution sources. Machine
learning and artificial intelligence are the emerging technologies which can identify
the trends and patterns and offers us an accurate and reliable prediction accuracy
[6]. Its ability to analyse large datasets, detect patterns and make prediction which
provide the advance data analysis for the problem statement. Traditional methods
fail to analyse all the factors because it is induced my multiple factors. It bridges
the gap by enabling the data-driven decision making and offering accurate
solutions.
1. Prediction and forecasting: One of the important applications of machine
learning is predicting about the pollutant and forecasting. Models can analyse
the historical data and make the prediction on that basis.
2. Source Identification: Algorithms are the master at identifying the patterns. By
providing the required data it can identify the major source from which air is
being polluted.

3
3. Real-Time Monitoring: This application allows to use the models in real-time
by providing them real-time data.

4
CHAPTER 2
LITERATURE REVIEW

2.1 Existing Research

Natarajan, Shanmurthy et. al. (2024) proposes optimized machine learning models
which combines grey wolf optimization (GWO) with Decision Tree (DT). Research on
major Indian cities demonstrated the effectiveness of GWO-DT model achieving in
higher accuracy. GWO-DT achieve better performance compared to other models of
91.49% for New Delhi[3].

Zayed and Abbod (2024) proposes a fuzzy logic framework to cover the blurry areas of
AQI prediction. Recent research integrated fuzzy air quality levels prediction (FAQLP)
with advance machine learning models. DNN-Markov model for accurate hourly AQI
prediction and adaptive neuro-fuzzy inference system (ANFIS). Root means square
error (RMSE) value is 15.27 for DNN-Markov model[6].

Binbusayyis, Khan et. al. (2024) proposed smart cities-based prediction model to utilize
the proposes regression algorithm. To achieve the model efficiency with pre-processing
of input data performed Deep Generative adversarial Network (GAN). Imputation,
feature scaling techniques have been used. Processed data, regression is done through
stacked attention GRU with KL divergence. R-Squared for the Deep-GAN is to be
0.94[4].

Liu, Cui et. al. (2024) focuses on integrating metrological factors with the data to
improve forecasting accuracy. Studies have identified metrological variables
influencing pollutant concentrations including temperature, humidity, air pressure etc.
Ten metrological factors were selected including seasonal characterises were analysed.
LightGBM achieved the higher performance 97.5% and LSTM model is accurate in
forecasting[7].

Karthick, Aruna et. al. (2024) focuses on predicting the AQI for Delhi using machine
learning techniques. The methodology encompasses steps such as data collection, pre-
processing, analysis and modelling. IterativeImputer were used for missing values
imputer. Box-Cox transformation was used for data normalization and variance.

5
Random Forest and XGBoost were the top performers based on model R-square 0.951
and 0.952[8].

Emeç and Yurtsever, [Link]. (2024) Explored the effectiveness of ensemble modelling for
predicting indicators like particulate matter 2.5 (PM2.5). Stacking ensemble model has
emerged by combining predictions from multiple machine learning algorithm. The
proposed stacking ensemble has performed well in compared to MLP, SVR, and RF
with r-square of 0.91, MAE 6.67 and RMSE 8.80[9].

Drewil , Al-Bahadili [Link]. (2024) proposes a deep citywide multisource data fusion-
based air quality estimation (FAIRY). Multisource data is considered and estimates the
air qualities of all region at a time. It generates the images form the citywide multisource
data that is traffic data, metrological data etc. and uses SegNet to learn the features from
these images. FAIRY refines low resolution fused features by employing high-resolution
fused features. It performs better and 15.7% effective than the best baseline[10].

Cao, Zhang, [Link]. (2024) proposes a hybrid model to simultaneously predict seven air
pollutants or the parameters from multiple monitoring stations. Model consist of three
components, Extended ARIMA to predicted the matrix series of multiple parameters
from multiple stations, EMD to decompose the air quality time series data into sub-
series and the truncated singular values decomposition (SVD) to compress and denoise
the expanded matrix. RMSE The proposed model is the lowest with RMSE of 12.10[11].

Sadriddin, Mekuria, [Link]. (2024) employed two methodologies that is random forest
regression and long-short term memory LSTM with the many metrological factors
including temperature and humidity. RFR outperformed as the superior predictor for
PM2.5 concentrations achieving accuracy of 95%[12]. Chatterjee, Raju et. al. (2024)
introduced Smart Cities Impact on Predictive Air Quality Management (SPAM) which
utilizes bidirectional stacking mechanism surpassing existing methods SVM, MLP,
RAQP, LSTM etc. [13]. Rad, Razmi et. al. (2024) employed XGBoost, LightGBM and
Random Forest to analyse CO, O3, NO2, SO2, PM10, PM2.5 concentrations. Random
Forest outperformed best for gaseous pollutants R-Squared CO=0.63, O3 =0.60, NO2 =
0.54 , SO2=0.55 while LGBM excelled in PM2.5 R-Squared of 0.88 and XGB in PM10
R-Squared of 0.88[14].

Elsarraj, Mahmoudi et. al. (2024) integrated Computational Fluid dynamics (CFD) with
a Eulerian-Lagrangian model to quantify infection risk using CO2 concentration and

6
ventilation effectiveness. Optimised Random Forest regression model achieved an R-
Squared of 0.99 and RMSE of 0,025 demonstrated higher accuracy and model
robustness[15].

Alolayan, Almutairi et. al. (2023) constructed three models using different
combinations. Model 1 consists of features related to the weather, Model 2 used only
features of the types and quantities of fossil fuels consumed by power plants, Model 3
include all the features from model 1 and model 2. The R-Squared of machine learning
increased on average using RF and SVM methods by 66% and 49% for NO2 models
and 27% and 23% for SO2 models[16].

Shao, Chen et. al. (2023) introduces a novel couped air quality optimization prediction
model based on Variational Model Decomposition (VMD), the informer time series
algorithm, Extreme Gradient Boosting (XGBoost) and the Dung Beetle Optimization
Algorithm (DBO). The proposed model achieved superior performance well in Nanjing
R-Squared of 0.96, RMSE of 1.98 and MAE of 1.624[17].

Berkani, Gryech et. al (2023) introduced a low-cost pollution monitoring system in


Morocco generating a unique air quality dataset. A comparative analysis of forecasting
models revealed that LightGBM and CatBoost out performed other approaches in short-
term PM pollutant predictions[18].

Iskandaryan, Ramos et. al. (2023) focused on spatiotemporal air quality prediction using
advance deep learning techniques. Attention Temporal graph Convolutional Network
A3T-GCN, combining Attention, GRU, GCN, demonstrated superior accuracy in
prediction. The Proposed model Outperformed with the RMSE of 19.14 and MAE of
15.33. Further studies aim to expand the geographical area, refine the graph structures
and optimize the model architecture for improved air forecasting [19].

Livingston, Kanmani et. al. (2023) briefs various state of art methods such as SVM, RF,
ANN, RNN and FL helps understanding the level of pollution. The results show
prominent outcomes which can be deployed in various cost-effective hardware
platforms for household and commercial purposes. The proposed model used eight
parameters like NO2CO, O3, PM2.5, PM10, SO2, TEMP, PRES, DEW, RAIN, WD,
WSPM[20]. F. Farhadi, R. Palacin et. al. (2023) explored the role of machine learning
in assessing air quality interventions by using machine learning models including
LGBM and LSTM for evaluating the impact of the clean air zones NO2 pollutant. Their

7
study demonstrates the effectiveness of the predictive models in validating policy
outcomes[21].

Al-Eidi, Amsaad et. al. (2023) has explored air quality prediction using machine
learning techniques. Random forest, Linear regression and Decision Tree were
compared for air quality forecasting. Decision tree outperformed mango the models
making it most effective due to high accuracy and low error rate. Methods like EDA and
SMOTER have enhance the accuracy of the decision tree model[22].

Alam, Hussain et. al. (2023) implemented various machine learning techniques like
CatBoost Regression, LightGBM Regression, XGBoost Regression, Extra Trees
Regression, Random Forest Regression, Artificial Neural Network, Support Vector
machine, Navie bayes and ARIMA. LightGBM performed well with accuracy about
0.997[23]. Han, Liu, [Link] (2023) proposes a self-supervised hierarchical graph neural
network (SSH-GNN) for fine grained air quality forecasting in semi-supervised way.
SSH-GGN first approximated the city-wides air quality distribution based on historical
readings and various contextual factors. Hierarchical recurrent graph neural network is
also proposed to make city-wide predictions which encodes the spatial hierarchy of
urban region for long range modelling. The proposed model has RMSE of 28.51 which
is very effective amongst for 1-6 hour [24].

Hossein , Motlagh [Link]. (2023) contributes a review of UAV-based air quality


monitoring, highlights and analyse the technical solutions and limitations and
identifying open challenges with the aim of providing a research roadmap. UAVs can
increase the spatial and temporal resolution of air quality data, especially by providing
insights into the vertical distribution of pollutants, facilitating searching and detecting
sources of emissions, and monitoring and auditing pollution distributions around fixed
sites, such as harbours or industrial plants [25].

Wardana, A. Fahmy, [Link]. (2023) Develops the low-cost air quality mentoring device
with embedded tinyML models. Two tinyML models were deployed for predicting air
quality and to recover the missing values. The coefficient ft the determination is
0.70[26].

Kalajdjieski, Trivodaliev et. al. (2023) proposes and evaluated four architectures of
encoder-decoder with attention for forecasting PM levels that are geographic a season

8
dependent. For handing missing values two adversarial networks are used for data
augmentation[27].

Dong, Zhang, [Link]. (2023) proposes a method based on empirical mode decomposition
(EMD), a transformer and a bidirectional LSTM which is good at ultrashort term
prediction of nonlinear time series data. The AQI sequence is first decompose intro
intrinsic mode functions (IMFs) via EMD. The Predicted values of each IMF are
integrated using BiLSTM to obtain the predicted AQI values. RMSE of the proposed
model for Patna is 5.68 [28].

Jin (2023) proposes a spatiotemporal graph neural network (BGGRU) a time series
prediction network with self-optimization for mining the pattern of time series. The
network consists of two module that is spatial and temporal modules. Graph sampling
and aggregation network (Graph SAGE) is used to extract spatial information and
Bayesian graph gated recurrent unit (BGraphGRU) is used for temporal module. The
accuracy comes out to be 99% for BGGRU proposed model[29].

Luo and Gong, (2023) proposes a method to accurately predict air pollutants with the
efficiency of the air pollution management. ARIMA is use to extract the linear part of
the air pollution data and output the nonlinear part, WOA-LSTM model is used to
predict the non-linear part. Whale Algorithm is used to adjust the hyper parameter for
the LSTM. The stability and accuracy of the proposed model is the highest with the R-
Squared of 0.97[30].

Zhang, Yang et. al. (2023) focuses on a robust forecasting system to achieve accurate
multi-step ahead forecasting for particulate matter (PM2.5 , PM10).Correlation analysis
is adopted to scree the spatial information. Spatial-Temporal attention mechanism is
used to assign weights to original inputs. This STA-ResCNN has RMSE of 6.986 which
is the highest amongst the compared models for the shanghai city[31].

Ravindiran, Hayder [Link]. (2023) Employed five machine learning models that is
XGBoost, Adaboost, Catboost, Random Forest, LightGBM. Catboost Model
outperformed other models with the r-square of 0.99. Adaboost was least effective
model with r-square of 0.97[32].

Mampitiya(2023) Focuses on the forecasting of the PM10 values in Sri Lanka. Five
Machine learning models were used that are extreme gradient Boosting (XGBoost) ,

9
LightGBM ,Catboost , GRU , LSTM with twelve air quality features in the dataset.
LightGBM performed better amongst with the R-square value of 0.99[33].

Kumari and Singh et. al. (2023) highlights the potential of the LSTM in providing the
forecasting of CO2 pollutant. The study explores multiple models for Co2 emission
prediction in India, including Autoregressive-integrated moving average (ARIMA),
seasonal ARIMA with exogenous factors (SARIMAX), Holt-winters, LR, RF and
LSTM. LSTM performs well compared to all the models with RMSE of 60.635 and
MAPE of 3.101% [34].

Maltare , Vahaora et. al. (2023) focuses on the support vector machine algorithm with
hyper parameter that is RBF kernel model. Studies have compared SARIMA, SVM and
LSTM models and highlighting their predictive capabilities for AQI forecasting. Data
Normalization, outlier removal and feature selection enhance the model quality. 0.96 R-
Squared was achieved with linear hyper parameter and 0.99 with RBF kernel[2].Zhao ,
Wu et. al. (2022) proposes a novel framework based on statistical learning integrating
SAC variables, feature selection and support vector regression for AQI prediction.
Forecasting performance is improved by techniques like feature selection, heuristic and
reinforcement learning[1].

Wang , Li et. al. (2022) presents a model based on convolutional neural network and
Improved Long Short-Term Memory CNN-ILSTM. ILSTM doesn’t have output gate
and input and forget gate is improved. ILSTM is efficient in learning historical data
which improve prediction accuracy. Eight prediction models were compared with
Shijiazhuang city air quality dataset[5].

Drewil , Al-Bhadili [Link] (2022) proposes a novel approach combining LSTM with the
genetic algorithm (GA) which works as an optimization technique to identify the best
hyper parameter for the model. It aims to predict the pollutants. This optimization results
in improved accuracy and efficiency in comparison to the traditional methods. RMSE
value for GA-LSTM is 9.5820 and MAE is 19.164[35].

Gilik , Ogrenci et. al. (2022) explored the univariate and multivariate approaches.
Univariate model will be focused on predicting concentration for one pollutant, where
the multivariate model will be focused on predicting multiple pollutants. It proposes a
novel method for air pollution by combing CNN and LSTM effectively capturing the
relationship for spatial-temporal data[36]. Heydari , Nezhad et. al. (2022) introduces a

10
novel hybrid model combining LSTM with mutli-verse optimization (MVO) algorithm.
The LSTM model is used as a forecasting engine while the MVO optimizes the
parameters. Among LSTM-PSO, ENN-MVO and ENN-PSO, LSTM-MVO is
performed better and leads to the better results[37].

Kim , Lim et. al. (2022) aimed to develop a short-term prediction model for PM10 and
PM2.5 using metrological data. Tree based Machine leaning algorithm is used and
performance was compared from the community multi-scale air quality (CMAQ)- based
chemical transport model (CTM). LightGBM has outperformed other tree-based ml
algorithms .CTM models can provide reliable and precise results by incorporating
ground and satellite-based observations. LightGBM efficiency and rapid learning
capabilities make it an idea decision[38].

Gladkova , Saychenko et. al. (2022) discusses the need to predict the changes in PM2.5
concentrations for air quality monitoring. Several machine learning models were used
that are ARIMA, Facebook prophet and LSTM. Detailed analysis and visualization of
the data is carried out. RMSE of ARIMA model is 12.45, Facebook Prophet Model is
12.24, LSTM model is 7.86[39].

Lin , Wang [Link]. (2022) proposes a new deep learning approach that integrate the spatial
and temporal correlation to detect the air anomalies. A weighted adjacency matrix is
established to characterize the spatial correlation and feature matrix is constructed to
characterize the temporal correlation. An Advance Deep learning technique is being
used that is Context augmented graph auto encoder (Con-GAE) [40]. Duan , Li et. al.
(2022) proposes a deep neural network-based approach which consist of a deep
distributed fusion network at station level short term prediction and a deep cascaded
fusion network for the city level long forecast. Data transformation pre-processing
technique the former network adopts a neural distributed architecture to fuse
heterogeneous urban data for capturing direct and indirect factors. The proposed model
achieves the high accuracy of 0.812[41].

Baldi , Delnevo et. al. (2022) presents a study in which predicting of indoor vehicle
environmental condition is done. Linear Regression, Random Forest, eXtreme Gradient
Boosting (XGB) and Multi-Layer Perceptron are used. Among these XGB performed
very effective with the accuracy of 0.97 for PM2.5 pollutant[42].

11
Wang , Jin et. al (2022) proposed a hybrid model based on CNN, AGU which can deal
with vanishing gradient and exploding gradient problems of RNN in air quality
forecasting. Data Adjustment Module (DAM) and attention mechanism is introduced to
AGU. The Proposed model achieved R-Squared of 0.957, MAE of 7.087 and MSE of
180.361 making it superior from others[43].

2.2 Research Gaps

The study's reliance on short-term seasonal forecasts is one of its main research
weaknesses. The model's suitability for long-term climate forecasting is still up for
debate, despite the fact that it accurately depicts temperature patterns and extreme
occurrences like heatwaves. Predictions that are limited to a single month may not
accurately reflect larger seasonal and annual trends because climate patterns change
throughout the year. A more thorough understanding of climate fluctuations would result
from expanding the dataset to encompass other seasons, such as winter, spring, and
autumn. Additionally, this would enhance the model's capacity to identify new trends
that are essential for climate adaptation plans, such slow warming or changing weather
patterns. Furthermore, the model may be better able to predict extreme weather
occurrences that do not closely follow seasonal cycles if the forecasting time is extended
beyond one month[44].

The absence of sophisticated data augmentation methods to enhance model performance


and stability represents another significant gap. AI models need a lot of training data to
generalize well, yet climate datasets are frequently small. The robustness of the model
may be improved by including data augmentation techniques as transfer learning from
other climatic datasets, perturbation-based augmentation, or synthetic data generation.
Predictive accuracy may be increased by using augmentation techniques such as
introducing noise into time-series data, fabricating fluctuations in meteorological
parameters, or utilizing generative adversarial networks. Furthermore, to integrate
various data sources and strengthen the model's resistance to overfitting, ensemble-
based augmentation techniques might be investigated. Researchers could create AI-
based climate models that function consistently across time periods and geographical
locations by filling this gap, producing more accurate climate projections[44].

In real-time air quality forecasting, managing noisy and missing data is a crucial issue
that has a big influence on model accuracy and dependability. Numerous data sources,

12
including as sensor networks, weather stations, and CCTV-based traffic surveillance,
are used by air quality monitoring systems. However, because of sensor failures,
network connectivity problems, severe weather, and data corruption, these sources
frequently have data gaps, transmission mistakes, and discrepancies. Unstable network
connections can result in delays or losses in data transmission, and sensor failures
brought on by power outages, hardware issues, or calibration mistakes might result in
missing data. Furthermore, climatic conditions like intense rain, fog, or storms can block
sensors and CCTV cameras, decreasing the amount of data available and adding noise
to the system[45][46].

Deep learning models have proven to be highly accurate at predicting intricate patterns
in a variety of domains, including air quality prediction. These models are ineffective
for real-time applications, nevertheless, because they frequently have higher computing
requirements. Deploying complex models on devices with low processing capacity, like
mobile devices, Internet of Things sensors, and edge computing systems, is challenging
due to their high computational requirements. Large models are also less effective in
time-sensitive applications like traffic management, pollution spike early warning
systems, or health warnings since they require more time to train and infer outcomes.
Optimizing model topologies to increase efficiency without sacrificing predictive
accuracy is becoming a crucial field of research as machine learning is further
incorporated into numerous real-world applications[47].

A crucial first step in guaranteeing high accuracy, effectiveness, and generalizability


across many datasets is optimizing a machine learning model. Numerous strategies are
used in model optimization, such as feature engineering, architectural enhancement,
hyperparameter tuning, and increases in computational efficiency. A model can preserve
practical usability in real-world applications while improving prediction performance
by carefully tuning these factors. Hyperparameter tweaking is one of the most important
parts of model optimization. Network designs, regularization intensities, and learning
rates are all governed by parameters in machine learning models. Methods like Bayesian
Optimization, Random Search, and Grid Search aid in methodically identifying the
optimal hyperparameters. In deep learning models, changing the number of layers,
neurons per layer, dropout rates, and batch sizes can greatly affect model
performance[48].

13
CHAPTER 3

PROBLEM FORMULATION AND PRESENT WORK

3.1 Problem Formulation

The lack of integration of external environmental variables is one of the biggest


obstacles in the prediction of the Air Quality Index (AQI). A complicated interaction
between natural events, human activity, and weather circumstances affects air pollution.
Unfortunately, a large number of current AQI prediction models ignore important
external environmental elements that have a substantial impact on air quality in favour
of relying solely on previous AQI data and a small set of air pollutant concentrations.
Temperature, humidity, wind speed, air pressure, and precipitation are some of the
climatic variables that directly affect how pollutants spread, change, and settle. For
instance, wind direction and speed affect how pollutants spread over different regions,
and high humidity levels can promote the development of secondary pollutants such
particulate matter (PM2.5 and PM10). Severe air pollution episodes can also result from
stagnant circumstances caused by changes in atmospheric pressure that trap pollutants
close to the surface. Conversely, precipitation aids in removing airborne contaminants,
momentarily enhancing the quality of the air.[49][50][51].
Air pollution levels are largely determined by human-induced activities, including
traffic density, industrial emissions, urbanization, and seasonal changes, in addition to
meteorological influences. While industrial operations contribute to sulphur dioxide
(SO₂) and volatile organic compounds (VOCs), which combine with sunlight to generate
ground-level ozone (O₃), traffic congestion, for example, increases emissions of
nitrogen oxides (NO₂) and particulate matter. Seasonal fluctuations make AQI
prediction even more difficult. While summer months may see increased ozone
generation due to brighter sunshine and higher temperatures, wintertime heating
emissions and temperature inversions can trigger severe pollution episodes[52][53].

When creating machine learning models, computational efficiency is crucial, especially


for real-time applications like autonomous systems, financial forecasting, and air quality
prediction. The computing requirements for training, inference, and storage rise in
tandem with the complexity of models in deep learning systems. Because of the
14
difficulties with scalability, real-time processing, and technology constraints, models
must be optimized for increased efficiency without sacrificing accuracy. Reducing
model complexity is a crucial component of computing efficiency. Because they have
so many parameters, many deep learning models, including CNNs and LSTMs, demand
a lot of processing power. Model size and inference time can be decreased by using
strategies like knowledge distillation, quantization, and pruning. Knowledge distillation
brings information from a bigger, more complex model to a smaller, quicker model;
quantization decreases the precision of model parameters (e.g., from 32-bit floating
point to 8-bit integers); and pruning eliminates superfluous weights from neural
networks[54].

Preprocessing guarantees that the data used for modelling is precise, consistent, and
appropriate for analysis, making it an essential stage in air quality prediction.
Inconsistencies in raw data obtained from monitoring stations are frequently caused by
a variety of variables, including human error, environmental issues, and sensor
malfunctions. To increase predictive models' overall performance and dependability,
these problems must be resolved. Managing incomplete or missing data is a crucial
component of preprocessing. Inaccurate projections may result from improperly
handled gaps in air quality datasets caused by transmission mistakes or device
breakdowns. Finding and adding the right values to these gaps enables the model to
produce predictions that are more informed by a full dataset. In a similar vein,
discrepancies in recorded values must be fixed to avoid deceptive data patterns. To
increase computational efficiency and model correctness, the dataset must be refined to
only contain the most pertinent information. Eliminating unnecessary or redundant data
makes the model simpler and guarantees that it concentrates on the elements that have
the biggest effects on air quality. To determine the most important criteria in predicting
air pollution levels, the interactions between different environmental aspects are also
investigated[55].

The capacity of sophisticated optimization techniques to effectively seek for ideal


solutions in high-dimensional areas is one of their main benefits. Conventional
optimization techniques frequently suffer from sluggish convergence, local minima, or
high computing expenses. In order to guarantee that models can find the optimal
parameter configurations for precise AQI prediction, advanced algorithms integrate

15
tactics that improve global search capabilities. By adjusting parameters, increasing
computational efficiency, and lowering mistakes, optimization methods are essential for
improving the performance of air quality forecast models. Advanced optimization
techniques aid in improving forecasting accuracy and robustness because predicting air
quality requires processing huge and complicated datasets. These techniques increase
the models' ability to adapt to changing environmental variables, avoid overfitting, and
improve model convergence. Furthermore, strategies to balance exploration and
exploitation are introduced by contemporary optimization approaches. By using
exploration, the algorithm is able to look widely across a variety of potential solutions,
avoiding an early convergence to less-than-ideal outcomes. By concentrating on
potential areas of the search space, exploitation, on the other hand, improves solutions.
For air quality models, which need precise parameters to take into consideration
fluctuating pollution levels and meteorological effects, this balance is essential[56].

3.2 Objectives

a. To pre-process the data for air quality index in India.

b. To design and implementation of hybrid machine learning model for prediction


of air quality index in India.

c. To compare the proposed hybrid machine learning model with existing state of
the art.

3.3 Proposed Methodology

The suggested methodology's typical steps are shown in this flowchart. Data
collection is the first step, when pertinent information is gathered from multiple
sources. The data is then cleansed and prepared for analysis during the next step,
data pre-processing. Model Design and Development involves the creation and
training of suitable models and algorithms. In order to guarantee accuracy and
dependability prior to deployment, Model Evaluation and Validation lastly
evaluates model performance.

16
Data Collection

Data Pre-processing

Model Design and Development

Model Evaluation and validation

Figure 3.1 Proposed methodology flow chart

Objective 1: To pre-process the data for air quality index in India.

3.3.1 Data Collection

Collection of data through Central Pollution Control Board (CPCB) is planned from
August 2024 to February 2025 for air quality prediction using machine learning in India.
Application Programming Interface (API) is used for collecting the dataset for this
period of time. Pollutants such as PM2.5, PM10, NO2 etc. including metrological
parameters like temperature, humidity etc were recorded. Time, Data and location stamp
were also recorded. These parameters contribute in identifying the air quality levels and
assessing their impact on the environment.

The data collection is automated using Python script in which Request is scheduled
every hour to capture real-time fluctuations in the air levels. The collected data were
stored in Supabase enabling efficient cloud-based management, easy accessibility, and
integration. To maintain the hourly request, cron job was setup, ensuring timely and
consistent updates.

3.3.2 Data preprocessing

An iterative process cycle for data preprocessing operations is depicted in this diagram.
Through repeated cycles of processes including data preparation, model training,
evaluation, and refinement, it places an emphasis on continual improvement. The
cyclical nature demonstrates that models are developed through trial and error and fine-
tuning to attain peak performance rather than being created in a single run.

17
Loading Dataset

Exploratory Data
Analysis

Handling Datatime
& Define feature
type

Imputation Process

Feature selection

Saving of
processed Data

Figure 3.2 Data Pre-Processing methodology flow chart

18
Loading Dataset: Using pandas Library, the Dataset stored in the Comma Separated
Values (CSV) format is read. For cleaning the dataset unnecessary columns were
removed to ensure proper structuring. Redundant columns are removed in this step to
avoid the confusion of the features. The Dataset consists 17 columns like city, state,
AQI, CO, dew, NO2, O3, p, PM2.5, PM10, SO2, t, w, etc., and 158455 rows.

Exploratory Data Analysis: It is the step to understand the data before applying the
machine learning or deep learning models. It involves the data summary, missing values,
relationships between features etc.

a) Missing values: It refers to the absent values in one or more feature or a row
b) Feature Distributions: The way that values of a specific feature are dispersed
throughout the dataset is referred to as the feature distribution. It helps us to
understand whether the data is skewed and has abnormalities or has a normal
distribution.

Handling Data-time & Define feature type: The dataset contains a column name
created_at, it is converted into a structured datetime format. Year, month, day, hour,
minute, second, and weekday were extracted from a single column.

Database consists of both numerical and categorical columns, Categorical features like
city, state and AQI_bucket contain textual or string values which need to handle
separately in this objective.

Imputation Process: In this objective two imputers are used to generate the missing
values:

a) K-Nearest Neighbours (KNN): KNN operates on the principle that similar


datapoints are often found close to each other in a feature representation. It is an
instance-based, non-parametric learning algorithm, which means it does not make
any assumptions about the underlying data distribution and relies on strong the
entire training dataset. A new datapoint needs to be classified or its value predicted
the algorithm identifies the “K” closest training samples to the input point using a
distance metric, typically Euclidean distance. Label among the K neighbours are
assigned to the new data point for classification tasks, Prediction is often the
average or the weighted average of the values of nearest neighbours for the
regression [Link] instances with similar known attributes are likely to have

19
similar missing attributes. The algorithm searches for the “K” most similar
instances in the dataset which are not missing for a specific feature. It finds similar
data points in its neighbours and fills the missing values based on the average of
the nearest neighbours (k). Number of nearest neighbours 5 is used (k=5) which
will estimate at the five most similar data points. KNNImputer from
[Link] is initialized with n_neighbors=5 and it will impute only numerical
features and will ignore categorical features[57].

b) Autoencoder-Decoder: It works by learning compress data into lower-


dimensional representation which is called latent space and then reconstruct it as
accurately as possible It consists of Encoder and decoder which are the main
components of this model. Encoder Takes the input and maps it to the compressed,
lower dimension form which is achieved using neural networks. The decoder takes
the compressed representation from the encoder and reconstructs the original
input. The goal is for the output of the decoder to be as close as possible to the
original input data. It is a type of ANN which is trained to re construct the dataset
by learning the Patterns in the existing data. But before directly applying the we
used SimpleImputer by mean method. Dataset is scaled using MinMaxScaler
which transforms the value between 0 to 1 Then Autoencoder refines the imputed
value by reconstructing the dataset with reduced errors. These techniques capture
the complex patterns in the data which is very robust. The Model consists of input
layer, several hidden layers and an output layer. Encoding process involves the
compressing the input into a lower -dimensional representation through fully
connect layers (128, 64, 32 neurons) with ReLU as activation function. The
Decoding process the then reconstructs the input data using symmetric layers (64,
128 neurons) and outputs the final prediction through sigmoid activation function.
Model is trained for 50 epochs with the batch size of 32[58].

Feature selection: it is performed to retain the relevant features and remove the
irrelevant features.

a) Correlation-based Feature selection: The Corelation Based feature selection


techniques is used to choose the pertinent features in a dataset by assessing the
correlation between and the target variable as well as the inter-correlation
between features. Features that have a strong correlation with the target but noit

20
with one another is the key objective to select the features. The premise of CFS
is that good feature subsets have feature that are not reductant to one another but
have a strong relationship to the output class. If two features have correlation
greater than 0.85, it means that provide similar information. One of the highly
correlated features is removed to avoid redundancy and reduce
multicollinearity[59].
b) Catboost Feature selection: It is a gradient boosting model optimized for
categorical data which is trained to determine the feature selection and
importance. It assesses each feature importance during training according to how
much it helps lower the model error. Features that result in splits that are more
accurate are prioritized. Feature with an importance score greater than 1% are
retained while less significant features are avoided. This step is crucial as it
contains only informative variables which improve accuracy of the model. It is
then trained on both datasets i.e. KNN-imputed and autoencoder-imputed
datasets. Columns like city , State , AQI_bucket contain categorical information
so these are passed first as cat_features , allowing model to treat them
appropriately during training. Once the model is trained on each dataset,
get_feature_importrance() is invoked which return the important features[60].

Objective 2: To design and implementation of hybrid machine learning model for


prediction of air quality index in India.

3.3.3 Model Design and Development

CatBoost: Categorical Boosting is an open-source toolkit with high performance and


gradient boosting which is developed for categorical data in regression and
classification. It belongs to Gradient Boosting Decision Tree (GBDT) algorithm family.
CatBoost follows the same ensemble learning methodology which is followed by GBDT
which is by successively constructing an ensemble of decision tree each of which is
training to reduce the residual of the ones that came before it, but in CatBoost overfitting
and target leakage is improved. It is a superior choice for tasks involving complicated
real-world datasets some it is good at handling categorical information.

Majority conventional machine learning models necessitate the numerical


representation of categorical features using techniques like label encoding or one-hot
encoding. These techniques result to high-dimensional data or introduce noise by

21
assigning arbitrary numerical values. CatBoost use a method called target-based
encoding, which substitute a statistical from the target variable. It captures the statistical
relationship between a category and the prediction target. Target leakage is one of the
crucial challenges for target -based encoding which causes overfitting when the model
incorrectly incorporates data from the target variable that wouldn’t be accessible during
real-world inference. To tackle this issue CatBoost uses a method which is called
ordered target statistics which maintain causality and remove leakage. In this approach
the dataset is randomly permuted during training and statistics are computed for each
data point Soley using the data points that came before it in that permutation.

Symmetric or oblivious decision tress is another important aspect of CatBoost. In


contrast to traditional decision trees, which aloe nodes to divide based on various
attributes and thresholds, CatBoost’s symmetric trees use the same splitting mechanism
for all branches at every level as a result, a balanced predictable tree structure with
excellent memory and computational efficiency. It improves the regularization which
keeps the model from overfitting and preserve performance on unknown data. This
structure allows for the faster inference since the decision path for any input is uniform
and can be easily optimized for parallel computing.

Ordered Boosting is a mechanism which is used by CatBoost which is to prevent


prediction shift which frequently causes overfitting in conventional gradient boosting
techniques. Every new tree in traditional gradient boosting is trained to reduce the
present model residuals. However, the same data is used to train the tree is typically
used to compute these residuals. CatBoost solves this problem by computing the
residuals without utilizing the future knowledge simulating a more realistic predicting
situation. CatBoost does this by repeatedly permuting the training data at random. It
determines the model forecast for each data point in a permutation based only on the
instance that precede it in the permutation. This indicates that the model solely uses
previous observations to calculate the residual, not any information from the target
value. This keeps the learning process casual and stops the model from picking up the
shortcuts from subsequent data.

Gated Recurrent Unit: GRU is an RNN version that keeps track of past inputs in order
to process sequential and time series data. It uses gating mechanism to manage the
information flow and decide what should be retained, updated, retained or forgotten over

22
the time. The Update gate and the reset gate which control how the hidden state is
updated at each time step are the two main gates that GRUs use. It is a powerful and
efficient neural network architecture for modelling sequential data. The Update and reset
gate provide a flexible way to manage the flow to information through time, making it
especially useful for tasks where both short- and long-term patterns need to be captured.
It is a perfect balance between model complexity and the performance, often being the
preferred choice when resources are limited.

Figure 3.3 Gated Recurrent Unit Architecture

The Hidden state, a vector symbolizes the networks memory and contains data from
earlier time steps, is at the core of the GRU. Gating mechanism that cleverly control the
flow of information over time, in contrast to traditional RNNs that just use the hidden
state and have trouble preserving long term information. Update Gate and Reset gate
are the two primary gates, GRU computes a candidate hidden state (ht) which represents
new information that might be added to the network memory. The combination of these
elements allows the GRU to adaptively capture dependencies over sort and long
sequences.

The reset gate is deigned to determine how much previous hidden state should be
ignored. This is crucial for allowing the GRU to selectively reset its memory when new
inputs that are significantly different from the past context.

𝑟𝑡 = 𝜎(𝑊𝑟 𝑥𝑡 + 𝑈𝑟 ℎ𝑡−1 + 𝑏𝑟 ) (1)

23
The update gate manages the balance between the previous hidden state and the new
candidate hidden state. It determines how much of the past information should be carried
forward into the current hidden state.

𝑧𝑡 = 𝜎(𝑊𝑧 𝑥𝑡 + 𝑈𝑧 ℎ𝑡−1 + 𝑏𝑧 ) (2)

The current input and a gated version of the prior hidden state which is impacted by
reset gate determine it. By introducing non -linearity and memory mixing this the
enables the network to efficiently synthesize both new and old data.

ℎ̅𝑡 = 𝑡𝑎𝑛ℎ(𝑊𝑥𝑡 + 𝑈(𝑟𝑡 ⊙ ℎ𝑡−1 ) + 𝑏) (3)

CatBoost-Gated Recurrent Unit:

In order to extract informative features, the pipeline starts with pre-processed data that
has been run through a CatBoost model. A GRU neural network made up of stacked
GRU layers with dropout for regularization and a final dense layer for output is then fed
these characteristics after they have been reconfigured for GRU input. The combination
improves the accuracy of final predictions by utilizing GRU's capacity to describe
temporal dependencies and CatBoost's ability to handle categorical variables.

Figure 3.4 Hybrid CatBoost-GRU general architecture

Pre-processed data from both the datasets is feed into CatBoost model. The training
process was conducted separately on two different datasets one through K-Nearest
Neighbours (KNN) and second for Autoencoder dataset. The model was built up with
2000 boosting iterations for each dataset, which allows the ensembled of trees to learn

24
intricate non-linear patterns in the data. To minimize the chance of overfitting and
guarantee steady convergence, a learning rate of 0.05 was used and a maximum tree
depth of 6. Early stopping mechanism was included to improve generalization even
further. In particularly the model tracked the loss on an independent validation set and
halted training after 25 rounds in which performance did not increase. This model
maintains its predictive ability by stopping at the ideal moment before overfitting to the
training set.

The prediction of CatBoost model were included for next step. The original scaled
feature set was concatenated with the CatBoost output. which is a condensed non-linear
interpretation of the input features. By proving the GRU network with both the raw
descriptives statistics and a higher-level abstraction that the tree-based model had learnt,
this fusion was intended to increase the representational capacity of the input data. In
order to adapt to the input format required by recurrent architectures like GRUs, the
enhanced feature vectors that were produced were subsequently reshaped into a 3-
dimensional array, with the temporal dimension specifically set to 1.

Considering its shown effectiveness in collecting temporal relationships and long-range


contextual information in time series and sequential data, the GRU was selected for this
phase with two GRU layers the architecture was intended to be stacked GRU network.
With the return_sequence=true set the first GRU layer, which consist of 64 hidden units,
could transfer a series of concealed stated to the subsequent recurrent layer. The model
was able to maintain temporal resolution and gradually improve feature representation
because to this design decision with 48 units and return_sequences=false the second
GRU layer compressed the output into a single context-aware vector that could be
utilized for regression.

Dropout layers were added after each GRU layer with dropout rate of 0.2 to enhance the
model’s generalization and lower the possibilities of overfitting instead of memorizing
the training data, this technique encourages the network to acquire robust parameters by
randomly deactivating a portion of neurons while training. The kernel weights of both
GRU layers weights of both GRU layers were subjected to L2 regularization with
lambda value as 0.005, By limiting the complexity of the model, this type of weight
decay penalizes high weight values and further deters overfitting.

25
To Prediction of the AQI value, a completely dense layer with single output neuron was
added at the last layer. The output layer was kept linear with no activation function
because the task is a regression task. The Mean Squared Error (MSE) loss function, a
popular option for regression tasks that penalizes greater errors more severely, was used
to create the model. For gradient-based optimization the Adam optimiser was used,
which combines the advantages of adaptive learning rate and momentum to boost
stability and speed up convergence. During training Mean Absolute Error (MAE)
Statistic was also monitored for performance. Early Stopping was used to guarantee
effective training and avoid overfitting. In order to preserve computing resources and
guarantee that optimize weights were restored, this callback tracked the model’s
validation loss and stopped training if no improvement was seen for five consecutive
epochs. Batch size of 64 was used to train each model for a maximum 50 epochs,
providing a balance between training efficient and learning stability. By Stimulating
complex temporal and sequential relationships in the supplemented dataset, this GRU-
based step was essential in improving the accuracy of AQI predictions.

Root Mean Square Error: RMSE is the square root of the average of the squared
differences between predicted and actual values. It signifies the standard deviation of
predicted errors in a single numerical value. It signifies the distance of errors between
predicted values and actual values from the dataset. It is widely used for measuring the
accuracy for the predictive models. It is specifically used in regression analysis as it
provides how well a model performed. This metric quantifies the average magnitude of
the model’s prediction error. Lower the value higher the accuracy.

It is a statistical tool to identify how these residuals are spread out. It demonstrates how
the data is allocated surround the best fit line.

1
𝑅𝑀𝑆𝐸 = √𝑛 ∑𝑛𝑖=1(𝑦𝑖 − 𝑦̂)
𝑖
2 (4)

Where:

 n = number of observations,
 𝑦𝑖 = actual observed value for the ith observation,

26
 𝑦̂𝑖 = predicted value for the ith observation,
 ∑𝑛𝑖=1 = summation over all observations.

R-Squared: R2 is known as coefficient determination which is a statistical method to


evaluate the goodness of fit of a regression model. It quantifies how well the
independent variables explain the variation in dependent variable. It measures the
proportion of the variance in the dependent variable that is predictable from the
independent variables in a regression model. Higher the value higher the model accuracy
or the performance. It defines how well a predicted values matches the original one.

It generates the value between 0 to 1. If the value is closer to 0 then the model performed
very low, if the value is near 1 then it has very high accuracy or performed well.

𝑆𝑆𝑟𝑒𝑠
𝑅2 = 1− (5)
𝑆𝑆𝑡𝑜𝑡

Where:
SSres = Residual Sum of Squares is the sum of the squared differences between actual
values (𝑦𝑖 ) and the predicted values (𝑦̂).
𝑖

𝑆𝑆𝑟𝑒𝑠 = ∑𝑛𝑖=1(𝑦𝑖 − 𝑦̂)


𝑖
2
(6)

SStot = Total sum of squares is the sum of the squared differences between the actual
observed values (𝑦𝑖 ) and the mean of the observed values (𝑦̅).
𝑆𝑆𝑡𝑜𝑡 = ∑𝑛𝑖=1(𝑦𝑖 − 𝑦̅)2 (7)

1
𝑦̅ = ∑𝑛𝑖=1 𝑦𝑖 (8)
𝑛

 n = number of observations,
 𝑦𝑖 = actual observed value for the ith observation,
 𝑦̂𝑖 = predicted value for the ith observation,
 𝑦̅ = mean of the actual observed values.

Mean Squared Error: MSE is metric that measures the average of the squared
differences between the actual and the predicted values. It provides the magnitude of
the error of the model. Smaller the value of MSE better the prediction accuracy. It is the
one of the most common metrics used to evaluate the performance of the predictive

27
models in regression tasks. It defined over range [0, +∞), where it is 0 when the
predicted value is exactly the same as the true value.

1
𝑀𝑆𝐸 = ∑𝑛𝑖=1(𝑦𝑖 − 𝑦̂)
𝑖
2
(9)
𝑛

Where:

 n = number of observations,
 𝑦𝑖 = actual observed value for the ith observation,
 𝑦̂𝑖 = predicted value for the ith observation,
 ∑𝑛𝑖=1 = summation over all observations.

Mean Absolute Error: MAE is a performance metric used the average magnitude of
error in predictions. It concentrates on the absolute difference between predicted and
actual values. It is one of the most commonly used metrics for evaluating. Lower the
value better the performance.
1
𝑀𝐴𝐸 = 𝑛 ∑𝑛𝑖=1 |𝑦𝑖 − 𝑦
̂𝑖 | (10)

Where:

 n = number of observations,
 𝑦𝑖 = actual observed value for the ith observation,
 𝑦̂𝑖 = predicted value for the ith observation,
 |. |= absolute value
 ∑𝑛𝑖=1 = summation over all observations.

CHAPTER 4

28
RESULTS AND DISCUSSION

4.1 Platform Used

Google Colaboratory (Google Colab) was used to carry out the full study and
implementation pipeline, which included data preparation, imputation (using KNN and
Autoencoder), feature selection (using CatBoost), and sophisticated regression
modeling (using CatBoost Regressor and GRU neural networks). Without the need for
expensive local hardware, Google Colab is a cloud-based platform that provides an
interactive Jupyter Notebook environment with free access to GPUs and TPUs, greatly
speeding up the training and assessment of deep learning models. Because of its ease of
use, adaptability, and capacity to manage huge datasets and intricate models in a
collaborative setting, this platform was selected.
Python 3.x served as the primary programming language for this project, utilizing its
extensive library and framework ecosystem for deep learning and machine learning
applications. Scikit-learn was used for imputation, scaling, and evaluation metrics, and
pandas, numpy, matplotlib, and seaborn were used for data processing and visualization.
Even though Google Colab was used for the majority of the trials, the project can also
be carried out locally with little gear. The following hardware requirements are
recommended for local runs: System requirements: Windows 7 or later; RAM: at least
4 GB (but 8 GB or more is advised for more seamless operation); Hard Drive: up to 1
TB, depending on the size of the dataset; and SSD: at least 256 GB to guarantee quicker
read/write speeds when dealing with large datasets and models.
Regarding software requirements, the project works with Python 3.x and needs a local
Jupyter Notebook or a development environment like Google Colab. TensorFlow/Keras
was used to implement the deep learning components, guaranteeing compatibility with
systems that are CPU and GPU based. All things considered, the open-source Python
ecosystem and Google Colab offered the perfect foundation for effectively creating,
evaluating, and improving the AQI prediction models.

4.2 Data Pre-processing Results

29
Different pollutants in the dataset have differing percentages of missing values; the
highest percentages of missing data are seen in PM10 (25.85%), SO2 (24.94%), and
O3 (22.90%).

Table 4.1 Missing values for all numerical feature

Pollutant Missing values Percentage


AQI 406 0.25
CO 32207 20.3
Dew 23066 14.5
H 1765 1.11
NO2 28928 18.25
O3 36299 22.90
P 16782 1.059
PM10 40963 25.85
PM2.5 1227 0.77
SO2 39523 24.94
T 1 0.000631
W 454 0.2865

However, there aren't many missing values for pollutants like T (temperature) and H
(humidity), suggesting that the recordings are comparatively complete. To preserve the
integrity of the analysis, effective imputation techniques are crucial, particularly for
contaminants with more than 20% missing data.

Figure 4.1 Missing values heat map for all numerical features

The dataset's distribution of missing values is shown visually in the heatmap. Lighter
streaks show missing records, while dark regions show complete data. Pollutants
including CO, NO2, O3, PM10, and SO2 show notable gaps, which is consistent with
the numerical summary's trend and emphasizes the necessity of reliable data imputation
methods.

30
Figure 4.2 Percentage of missing values

The percentage of missing values for each pollutant is shown in the bar chart. With more
than 22% of the data missing, PM10, SO2, and O3 had the most gaps in the data.
Features like T, AQI, and W, on the other hand, exhibit little missing, indicating that the
recordings for those variables are trustworthy.

Table 4.2 Identification of categorical features

Categorical features
City
State
AQI_bucket

Table 4.3 Identification of Numerical features

Numerical Features
id

31
aqi
CO
dew
h
NO2
O3
p
PM10
PM2.5
SO2
t
w

The dataset consists of both categorical and numerical features. Categorical features
include City, State, and AQI_Bucket, which represent location and air quality
classification. Numerical features such as AQI, CO, NO2, PM10, and SO2 represent
pollutant concentrations and environmental measurements, essential for analyzing and
predicting air quality trends.

Figure 4.3 Feature Distributions all over the features

The dataset's numerical feature distribution is shown by the histograms. The majority of
variables, including CO, NO2, O3, PM10, and SO2, have distributions that are right-
skewed, meaning that there are few severe outliers and a preponderance of lower values.

32
Features that exhibit more bell-shaped or regularly distributed patterns are Dew,
Pressure (p), and Temperature (t). Understanding the magnitude and skewness of the
data is made easier with the aid of this visualization, which is essential for modeling and
preprocessing.

Figure 4.4 Correlation Heatmap of the features

It is generated to visualise the correlation between numerical features. This helps in


identifying highly correlated variables that may or may not be redundant. The linear
relationships between numerical features are displayed in the correlation heatmap. There
are significant positive relationships found between temperature (t) and wind speed (w),
as well as between PM10 and PM2.5. On the other hand, there is a significant negative
association between AQI and dew, suggesting that improved air quality may be linked
to greater dew levels.

33
Figure 4.5 Pair plots of features

The pair plot illustrates the connections between important air quality metrics, including
AQI, PM10, PM2.5, NO2, O3, and SO2. There are clear positive linear trends between
PM2.5 and PM10, indicating that they frequently rise simultaneously. There are non-
linear connections and possible outliers in other combinations, like AQI vs.
PM2.5/PM10, which exhibit increasing trends with discernible spread. Finding
associated variables and patterns that are helpful for model development is made easier
with the aid of this representation.

34
Figure 4.6 AQI Patterns over time

With multiple abrupt peaks signifying pollution spikes, the line plot depicting AQI
trends over time shows an overall increase from August 2024 to January 2025. The
variability or confidence interval, which highlights times when air quality fluctuates
more, is represented by the shaded region surrounding the line. Seasonal variations or
higher emissions in some months could be the cause of this rising trend.

Figure 4.7 Feature Importance using both feature selection techniques for KNN
imputation dataset

Based on KNN-imputed data, the feature importance plot demonstrates that AQI_bucket
is by far the most significant predictor. The effectiveness of the model is also greatly
influenced by environmental variables including temperature, CO, NO2, and dew point.
The fact that temporal characteristics like minute, second, and weekday have no effect
suggests that they have little predictive power in this situation.

35
Figure 4.8 Feature Importance using both feature selection techniques for Autoencoder
imputation dataset

According to the Autoencoder-imputed data feature importance plot, AQI_bucket is the


most important predictor, followed by CO, O₃, and NO₂. The quality of the air is
significantly influenced by these pollutants. While temporal and location-based features
like Year, Hour, and City have little effect on model predictions, environmental elements
like pressure (p) and wind speed (w) also have a moderate impact.

36
Table 4.4 KNN Imputed dataset feature importance

[Link] Feature Name Importance Value Importance percentage


1 AQI_bucket 64.294188 64.29
2 Dew 8.1189 8.12
3 no2 5.0034 5.00
4 T 4.3755 4.38
5 state 3.8643 3.86
6 Co 3.6119 3.61
7 H 2.2538 2.25
8 P 1.7474 1.75
9 id 1.5464 1.55
10 city 1.4538 1.45
11 o3 1.1337 1.13
12 w 1.0630 1.06
13 so2 0.9464 0.95
14 day 0.2269 0.23
15 Hour 0.1549 0.15
16 Weekday 0.1388 0.14
17 Second 0.0659 0.07
18 Year 0.0000 0.00
19 minute 0.0000 0.00

The feature importance values for the KNN imputed dataset. The AQI_bucket feature
has the highest importance at 64.29%, indicating it plays a dominant role in the model,
followed by Dew and NO2 with significantly lower importance percentages. Several
temporal features like Year and Minute have negligible or zero importance in this
analysis.

37
Table 4.5 Autoencoder Imputed dataset feature importance

[Link] Feature Name Importance Value Importance percentage


1 AQI_bucket 58.143000 58.14
2 no2 8.624802 8.62
3 Co 8.459964 8.46
4 State 7.115321 7.12
5 o3 3.567186 3.57
6 dew 3.458185 3.46
7 W 2.165282 2.17
8 P 2.146748 2.15
9 Minute 1.499414 1.50
10 city 1.148745 1.15
11 T 0.712388 0.71
12 Id 0.694390 0.69
13 H 0.678911 0.68
14 so2 0.484676 0.48
15 Weekday 0.368888 0.37
16 Day 0.266432 0.27
17 Second 0.203728 0.20
18 Year 0.148150 0.15
19 Hour 0.113789 0.11

The feature importance for the Autoencoder imputed dataset. Similar to the KNN
imputation, AQI_bucket is the most significant feature, contributing 58.14% to the
model. Other important features include NO2, CO, and State, each with importance
values between 7% and 8.6%, while temporal features like Year and Hour have much
lower importance.

After applying both features selection techniques to both datasets we get to know that
AQI_bucket emerged as the most important feature, holding the highest prediction
potential. Its relative importance does change substantially though. AQI_bucket makes
up over 65% of the total feature significance in the KNN imputed dataset, indicating
that the model mostly relies on this feature. After this feature dew, NO2, t, state, CO, h,
p, city etc.

However, AQI_bucket still holds the first priority in the autoencoder-imputed dataset,
although its significance is 55%. CO, O3, NO2, p can now make a more significant

38
contribution. Its is to be noted that time-related variables like minute were almost
insignificant in the KNN-imputed dataset but value is increased in Autoencoder Imputed
dataset. In KNN-Imputed dataset after AQI_bucket dew and in autoencoder imputed
dataset NO2 is the second feature which is importance. So, both dataset feature
importance is not same which will create a difference in evaluation metrics which we
will analyse in further sections.

Figure 4.9 Feature Importance Comparison

The heatmap illustrates how the relative value of various features in KNN and
autoencoder-based imputation varies. In the Autoencoder approach, for instance, NO₂
and CO are more significant, whereas in KNN, dew and temperature (t) are more
significant. This demonstrates how the features are prioritized differently by each
approach during prediction.

39
Figure 4.10 CatBoost-GRU Prediction VS True AQI on both datasets

The scatter plots provide a visual comparison of the predictive performance of the
CatBoost-GRU hybrid model utilizing two distinct imputation techniques. The overall
findings make it abundantly evident that although both models show a high degree of
positive correlation between the actual and predicted AQI values, the model trained on
autoencoder based dataset performs better than the KNN-based dataset. Better accuracy
and consistency are indicated by the autoencoder figure which shows tighter clustering
of predictions along the diagonal line, prearticular across low to middle AQI ranges. On
the other hand, especially at higher AQI values, the KNN-based model exhibits more un
predictability and dispersed predictions., indicating possible limits in the feature set’s
capacity to generalize to more severe circumstances. The Catboost-GRU can produce
more accurate predictions because the autoencoder based dataset selected features
capture more abstract and pertinent patterns form the original datasets.

Figure 4.11 CatBoost-GRU Prediction VS True AQI on both datasets

40
In the KNN dataset the predicted values roughly flow the general trend of the true AQI
values, there are evident discrepancies, especially around sharp spikes and high AQI
instances. This indicates that the model trained on KNN-selected features may struggle
to fully capture the complex dependences in the data. Where the autoencoder dataset
shows much closer alignment between the predicted and true AQI values, with the
predicted curve almost mirroring the actual values across the sample indices. Overall,
the Autoencoder dataset is producing more accurate results.

Figure 4.12 CatBoost-GRU Residual lot on both datasets

The residuals exhibit a clear funnel-shaped pattern indicating heteroscedasticity where


prediction errors increase with the higher predicted AQI values for the KNN dataset.
Autoencoder dataset more randomly scattered distribution of residuals around the zero
line, implying the errors are more uniformly distributed and model maintains consistent
performance across different AQI range.

41
4.3 Model Design and Development Results

Figure 4.13 CatBoost-GRU Model RMSE comparison on both datasets

RMSE values for CatBoost-GRU models employing KNN and Autoencoder imputed
data are compared in the bar chart. In comparison to KNN , the Autoencoder technique
produces a lower RMSE , suggesting superior predictive accuracy.

Figure 4.14 CatBoost-GRU Model R2 comparison on both datasets

Using CatBoost-GRU, the R2 comparison table demonstrates that both KNN and
Autoencoder imputation techniques produce good model correctness. Autoencoder, on

42
the other hand, performs somewhat better, obtaining a little higher R2 score, which
suggests better model fit.

Figure 4.15 CatBoost-GRU Model MSE comparison on both datasets

The CatBoost-GRU model trained on autoencoder selected features produces a small


Mean squared error than the model trained on the KNN-selected features, according to
the MSE comparison table. This implies that the Autoencoder imputed method produced
more accurate AQI predictions by offering more efficient feature representations. Better
performance and less prediction error are shown by lower MSE values.

43
Figure 4.16 CatBoost-GRU Model MAE comparison on both datasets

The KNN and Autoencoder feature selection techniques produced comparable mean
absolute error levels when used to the CatBoost-GRU model, as demonstrated by the
MAE comparison table. On other hand the autoencoder-based model’s average
prediction errors are significantly larger as indicated by its slightly higher MAE. In spite
of this the difference is negligible indicating that both strategies are equally successful
in reducing absolute forecast deviations.

Table 4.6 Evaluation metrics comparison on both dataset

CatBoost-GRU Model
KNN Dataset Autoencoder Dataset
RMSE 23.8605 RMSE 20.7190
MSE 569.3227 MSE 429.2787
MAE 11.4761 MAE 11.7035
R2 0.9785 R2 0.9838

44
Figure 4.17 CatBoost-GRU Model evaluation metrics on both datasets

Applying CatBoost-GRU Hybrid model to both the KNN-based and Autoencoder based
feature selected datasets for AQI prediction consistently shows high predictive
performance with excellent R-square scores 0.9785 for the KNN dataset and an even
higher 0.9838 for the autoencoder dataset indicating that the model is able to capture
and generalize the underlying patterns in air quality data , with R- square values showing
that a very high proportions of the variance in AQI values is explained by the model
predictions. However, a closer look at the performance metrics shows that the model
performs noticeably better on the Autoencoder-based datasets.

The Mean square Error MSE decreases dramatically from 569.3227 to 429.2787 while
the Root Mean Squared Error (RMSE) decreases from 23.8605 on the KNN dataset to
20.7190 on the Autoencoder dataset. These reduces error values suggest that, when
employing features acquired from the autoencoder technique the model predictions are
not only more accurate but also more consistent and closely clustered around the
genuine AQI Values.

Interestingly, the KNN dataset has a little superior Mean Absolute Error (MAE) meaning
that the average deviation is slightly lower in absolutes terms. The more significant gains
in RMSE, MSE and R-square are not offset by this small change though. Because RMSE

45
penalizes greater deviations more severely than MAE, which handles all mistakes
equally and is less sensitive to large errors, there may be a little disparity. The
Autoencoder dataset’s overall improvement in RMSE and R-square indicated improved
generalization and robustness, particularly in capturing more intricate, non-linear
relationships between the input features and AQI values.

These findings demonstrate the benefits of combining a sequential model like GRU
which is skilled at capturing temporal dependences, with a gradient boosting decision
tree model like CatBoost which is excellent at handling categorical data and learning
complex feature relationships. The hybrid CatBoost-GRU model can produce
predictions that are extremely accurate and stable when combined with a rich feature
space that is obtained by Autoencoder based dimensionality reduction. The enhanced
autoencoder dataset performance suggest that this feature selection technique is more
successful at preserving pertinent information while removing noise and redundancy,
which improves downstream model’s ability to learn. Using CatBoost-GRU hybrid
architecture, it is clear from this examination that the autoencoder-based feature selected
dataset offers a more appropriate and effective representation for AQI prediction.

Figure 4.18 Prediction Distribution on CatBoost-GRU model

Clear visual insights into the CatBoost-GRU hybrid model performance are shown buy
the LKDE plots that contrast the true and predicted AQI values for the KN-based dataset
and Autoencoder -based dataset. The projected distribution in the KNN-based dataset
mostly resembles the actual AQI distribution especially in the moderate AQI range of
roughly 50 to 250.A wider spread in the prediction curve and minor peak misalignments,
however imply that the model adds more variance and is less accurate in capturing the
true distribution. Both the true and projected curves show a noticeable increase around
AQI levels close to 1000 suggesting that the model can detect severe AQI cases.

46
4.4 State of the Art Methods Results

Table 4.7 Comparison table of all models on Autoencoder dataset


Autoencoder dataset
CNN- LSTM
CatBoost-GRU ILSTM[5] CatBoost-TCN LSTM[36] CNN-GRU [33]
RMSE 20.9267 30.3574 20.745188 35.8272 25.943 36.3154
MSE 437.9258 921.5698 430.362844 1283.5874 673.0413 1318.8106
MAE 11.3266 15.7635 11.663625 15.0017 12.1955 17.4577
R2 0.9835 0.9652 0.983745 0.9495 0.9746 0.9502
Time(sec) 365.2 840.68 3446.23 533.60 3314.83 865.72
Expense 31,057 35,969 93,377 51,457 88,001 12,051

Figure 4.19 Comparison of Evaluation Metrics for all models on KNN dataset

47
Table 4.8 Comparison of all models on KNN dataset
KNN Dataset
CatBoost-GRU ILSTM CatBoost-TCN CNN-LSTM CNN-GRU LSTM
RMSE 24.1831 27.8658 26.095111 28.8951 25.4571 34.5155
MSE 584.8221 776.5044 680.954797 834.9257 648.0661 1191.3217
MAE 11.6775 13.9582 14.227012 13.418 11.7063 17.5151
R2 0.978 0.9707 0.974342 0.9672 0.9756 0.9551
Time(sec) 425.72 906.83 3231.01 708.44 3112.43 888.23
Expense 31,633 36,737 93,889 52,033 88,001 12,651

Figure 4.20 Comparison of Evaluation Metrics for all models on Autoencoder dataset

48
CatBoost-GRU vs ILSTM

Table 4.9 Comparison of both dataset for CatBoost-GRU and ILSTM


Autoencoder dataset KNN Dataset
CatBoost-GRU ILSTM CatBoost-GRU ILSTM
RMSE 20.9267 30.3574 24.1831 27.8658
MSE 437.9258 921.5698 584.8221 776.5044
MAE 11.3266 15.7635 11.6775 13.9582
R2 0.9835 0.9652 0.978 0.9707
Time(sec) 365.2 840.68 425.72 906.83
Expense 31,057 35,969 31,633 36,737

Figure 4.21 R-Square comparison for CatBoost-GRU vs ILSTM

CatBoost-GRU continues to outperform ILSTM on all performance metrics in the KNN-


based feature selected dataset, following the pattern seen in the Autoencoder dataset. On
the KNN dataset, CatBoost-GRU’s RMSE is 243.18 whereas ILSTM records a larger
RMSE of 27.87, demonstrating once more how much CatBoost-GRU’s prediction differ
from the actual AQI values. Particularly when tracking pollution levels that affect public
health, even modest increases in RMSE can yield substantial practical benefits in
regression issues like AQI forecasting. Compared to ILSTM, which scores 13.96 on
MAE CatBoost-GRU scores 11.68. with a lower average error and better daily air

49
quality predictions, these figures show that CatBoost-GRU keeps producing forecasts
that are increasingly accurate and consistent.

The R2 score, which is 0.978 compared to ILSTM’s 0.9707 further demonstrates


CatBoost-GRU’s superiority CatBoost-GRU once again exhibits a better fit to the
underlying data distribution, indicating a stronger capture of complex, nonlinear
relationships. This is particularly advantageous in environmental datasets that are
influenced by numerous correlated variables , even though both models perform well in
explaining variance. In terms of training efficiency CatBoost-GRU completes training
far quicker than ILSTM, taking 425.72 seconds as opposed to 906.83 seconds. Ther
shorter training time allows for frequent model retraining, which improves adaptability
in changing situations and speeds up deployment cycles. Additionally, CatBoost-GRU
has a lower computational cost then ILSTM, costing 31,633 units as opposed to 36,737
units. This represents a significant 14% cost decrease that can scale significantly in
systems that are cloud-based or edge-deployed.

Overall, CatBoost-GRU offers superior performance, efficiency, and affordability on the


KNN dataset as well. ILSTM remains a valid choice for simpler modelling pipelines or
legacy systems but cannot match CatBoost-GRU's performance and scalability,
especially for high-precision forecasting applications.

CatBoost-GRU VS CatBoost- TCN

Table 4.10 Comparison of both dataset for CatBoost-GRU and CatBoost-TCN


Autoencoder dataset KNN Dataset
CatBoost-GRU CatBoost-TCN CatBoost-GRU CatBoost-TCN
RMSE 20.9267 20.745188 24.1831 26.095111
MSE 437.9258 430.362844 584.8221 680.954797
MAE 11.3266 11.663625 11.6775 14.227012
R2 0.9835 0.983745 0.978 0.974342
Time(sec) 365.2 3446.23 425.72 3231.01
Expense 31,057 93,377 31,633 93,889

50
Figure 4.22 R-Square comparison for CatBoost-GRU vs CatBoost-TCN

When comparing CatBoost-GRU and CatBoost-TCN across both the Autoencoder and
KNN feature-selected datasets, we observe that while both models are high-performing
and capable of accurate AQI prediction, CatBoost-GRU consistently offers a more
favourable balance between performance, training efficiency, and computational cost.
On the Autoencoder dataset, CatBoost-TCN shows a marginal edge in some error
metrics, with a slightly lower RMSE (20.75 vs 20.93) and MSE (430.36 vs 437.93), as
well as a virtually identical R² score (0.9837 vs 0.9835). However, CatBoost-GRU
compensates with a slightly better MAE (11.33 vs 11.66), indicating it may make fewer
large prediction errors. Despite this narrow difference in accuracy, CatBoost-TCN’s
training time of over 3446 seconds and cost of 93,377 units makes it vastly more
resource-intensive than CatBoost-GRU, which trains in just 365 seconds and costs only
31,057 units highlighting CatBoost-GRU’s efficiency advantage.

The contrast becomes even more pronounced on the KNN dataset. Here, CatBoost-GRU
not only remains more efficient but also surpasses CatBoost-TCN in all major accuracy
metrics, achieving a lower RMSE (24.18 vs 26.10), MSE (584.82 vs 680.95), and MAE
(11.68 vs 14.23), alongside a higher R² score (0.978 vs 0.9743). Additionally, it
maintains a significantly shorter training time (425.72 seconds vs 3231.01 seconds) and
lower expense (31,633 vs 93,889 units). This demonstrates that while CatBoost-TCN
can perform well, its computational overhead far outweighs its marginal or non-existent

51
gains in predictive accuracy, especially in more compact feature sets like those from
KNN selection.

In conclusion, while CatBoost-TCN can be slightly more accurate under specific


conditions, CatBoost-GRU proves to be the more versatile, efficient, and cost-effective
model across both datasets. Its consistent performance and lower computational
demands make it a highly scalable solution for real-time or large-scale AQI forecasting
applications, offering near-best-in-class accuracy without the burden of heavy training
time or cost.

CatBoost-GRU VS CNN-LSTM

Table 4.11 Comparison of both dataset for CatBoost-GRU and CNN-LSTM


Autoencoder dataset KNN Dataset
CatBoost-GRU CNN-LSTM CatBoost-GRU CNN-LSTM
RMSE 20.9267 35.8272 24.1831 28.8951
MSE 437.9258 1283.5874 584.8221 834.9257
MAE 11.3266 15.0017 11.6775 13.418
R2 0.9835 0.9495 0.978 0.9672
Time(sec) 365.2 533.60 425.72 708.44
Expense 31,057 51,457 31,633 52,033

Figure 4.23 R-Square comparison for CatBoost-GRU vs CNN-LSTM

52
In terms odd predictive accuracy, training efficiency and computational cost CatBoost-
GRU is unquestionably the better model when compared to CNN-LSTM on both
datasets. CatBoost-GRU performs better than CNN-LSTM on the Autoencoder dataset
in every important assessment metric. Given a substantially lower RMSE 20.93 vs
35.83, it shows a markedly decreased prediction error. This difference 437.93 vs
1283.59 is further highlighted by the MSE comparison which indicates that CatBoost-
GRU’s predictions are not only more accurate but also more reliable. It also makes fewer
big mistakes, as evidenced by its lower MAE 11.33 vs 15.00. Crucially CatBoost-GRU’s
R2 score 0.9835 is significantly higher than CNN-LSTM’s 0.9495 indicating that
CatBoost-GRU is better at capturing the underlying data variation. In terms of
performance CatBoost-GRU is significantly more effective for real world
implementation with a significant advantage in training time 365.2 vs 533.6 seconds and
computational cost 31,057 vs 51,457

The pattern is maintained on the KNN dataset, where CatBoost-GRU outperforms


CNN-LSTM in every important metric. Its lower RMSE (24.18 vs 28.90), MSE (584.82
vs 834.93), and MAE (11.68 vs 13.42) demonstrate its better forecasting accuracy and
stronger generalization. This is corroborated by R2 which shows that CatBoost-GRU
scored 0.978 and CNN-LSTM scored 0.9672. CNN-LSTM still performs worse than its
Autoencoder-based counterpart, despite improvements in accuracy and efficiency. For
resources constraints or real-time contexts, CatBoost-GRU continues to be substantially
faster to train (425.72 seconds vs 708.44 seconds) and more computationally
inexpensive (31,633 vs 52, 033).

CatBoost-GRU VS CNN-GRU

Table 4.12 Comparison of both dataset for CatBoost-GRU and CNN-GRU


Autoencoder dataset KNN Dataset
CatBoost-GRU CNN-GRU CatBoost-GRU CNN-GRU
RMSE 20.9267 25.943 24.1831 25.4571
MSE 437.9258 673.0413 584.8221 648.0661
MAE 11.3266 12.1955 11.6775 11.7063
R2 0.9835 0.9746 0.978 0.9756
Time(sec) 365.2 3314.83 425.72 3112.43
Expense 31,057 88,001 31,633 88,001

53
Figure 4.24 R-Square comparison for CatBoost-GRU vs CNN-GRU

Across both datasets, it is evident that CatBoost-GRU consistency performs better


overall than CNN-GRU, particularly wen it comes to accuracy and computational
efficiency. RMSE of 20.93 CatBoost-GRU shows significantly higher predictive
accuracy on the Autoencoder dataset Than CNN-GRU, which produces bigger
prediction errors on average. This difference is also observed in MAE and MSE
demonstrating that CatBoost-GRU generates fewer and milder outliers in its predictions.
Moreover CatBoost-GRU achieved a higher R2 score of 0.9835 than CNN-GRU which
indicates that it better captures underlying patterns in the data and explains a larger
portion of the variance in AQI levels. Compared to CNN-GRU which takes 3314.83
seconds to train, CatBoost-GRU is significantly faster from an efficiency perspective,
taking just 365.2 seconds. It also significantly less expensive computationally (31,057
vs 88,001 units) which makes it perfect for frequent retraining or development in time
sensitive scenarios

The KNN dataset shows the same performance disparity RMSE is 243.18 over CNN-
GRU’s 25.46, MSE is 584.82 versus 648.07, and MA is 11.68 versus 11.71 which is
very narrow difference in MAE , but CatBoost-GRU Continues to outperform in all key
accuracy parameters. Although both models are rather good at identifying data patterns,
CatBoost-GRU even has a higher R2 Score. But the resource efficiency is what makes
them stand out CatBoost-GRU trains far faster than CNN-GRU (425.72seconds versus

54
3112.43) These distinctions are crucial for situations including resource constraints,
frequent retraining on streaming data, and scaling up model deployment.

In summary, across both datasets, CatBoost-GRU not only delivers slightly better or
comparable prediction accuracy to CNN-GRU but also vastly outperforms it in training
speed and computational cost. While CNN-GRU performs reasonably well and benefits
from the strengths of convolutional and recurrent layers, the hybrid integration of
CatBoost with GRU appears to more effectively leverage structured features and
temporal dynamics in a lighter, more optimized framework making it the more practical
and powerful choice for AQI forecasting.

CatBoost-GRU VS LSTM

Table 4.13 Comparison of both dataset for CatBoost-GRU and LSTM


Autoencoder dataset KNN Dataset
CatBoost-GRU LSTM CatBoost-GRU LSTM
RMSE 20.9267 36.3154 24.1831 34.5155
MSE 437.9258 1318.8106 584.8221 1191.3217
MAE 11.3266 17.4577 11.6775 17.5151
R2 0.9835 0.9502 0.978 0.9551
Time(sec) 365.2 865.72 425.72 888.23
Expense 31,057 12,051 31,633 12,651

Figure 4.25 R-Square comparison for CatBoost-GRU vs LSTM

55
The performance differences between CatBoost-GRU and the standalone LSTM model
are both pronounced and persistent across both datasets, demonstrating CatBoost-GRU's
superiority in almost every way. With a substantially lower RMSE of 20.93 on the
Autoencoder dataset than LSTM's much larger value of 36.32, CatBoost-GRU
demonstrates its capacity for more accurate prediction. This disparity is further shown
by the MSE (437.93 vs. 1318.81) and MAE (11.33 vs. 17.46), where CatBoost-GRU
provides a more stable and broadly applicable model in addition to generating less
significant mistakes. CatBoost-GRU is significantly better at explaining the variance in
AQI levels and identifying significant temporal and contextual patterns in the data, as
seen by the significant difference in R2 scores between it and LSTM (0.9835) and
CatBoost-GRU (0.9835). Computational efficiency is a significant strength in addition
to accuracy: While LSTM takes 865.72 seconds, more than twice as long, CatBoost-
GRU trains in just 365.2 seconds. Furthermore, despite LSTM's simplicity, CatBoost-
GRU is more cost-effective, costing just 31,057 units as opposed to 12,051 units,
demonstrating that CatBoost-GRU provides higher returns for the money.

KNN dataset, this pattern persists. All measures, including RMSE (24.18 vs. 34.52),
MSE (584.82 vs. 1191.32), and MAE (11.68 vs. 17.52), show improved accuracy with
CatBoost-GRU, and the R2 value is greater (0.978 vs. 0.9551). These variations
demonstrate how much more effective the CatBoost-GRU hybrid architecture is than a
pure LSTM, particularly when working with reduced or optimized feature sets like those
in KNN. CatBoost is used to extract deep feature representations, while GRU is used to
model temporal dependencies. While LSTM has a lower computational cost (12,651
units), the decrease in predictive power renders these savings less significant in practice,
particularly when model accuracy is crucial for alerts or decision-making. In terms of
training time, CatBoost-GRU again outperforms 888.23 seconds (425.72 seconds).

56
CHAPTER 5
CONCLUSION AND FUTURESCOPE

5.1 Conclusion

Air quality refers to the health or condition of the air in our environment, which can be
adversely affected by the concentration of pollutants and the presence of harmful
substances. Predicting air quality is essential for safeguarding public health and
supporting environmental management. This study evaluates several machine learning
models for air quality prediction using two datasets Autoencoder and KNN reveals
significant performance differences in terms of accuracy, computational efficiency, and
cost. Among the tested models, CatBoost-GRU and CatBoost-TCN consistently
achieved the lowest RMSE, MSE, and MAE values, along with high R² scores (above
0.97), indicating strong predictive accuracy. ILSTM and CNN-based models showed
moderate performance, while the standalone LSTM model had the weakest results
across both datasets. Although CatBoost-TCN offered high accuracy, it incurred
substantial computational time and cost, making CatBoost-GRU a more balanced option
with competitive accuracy and lower resource demands. These findings suggest that
hybrid models combining tree-based algorithms with deep learning, particularly
CatBoost-GRU, provide a robust and efficient solution for air quality forecasting.
Overall, machine learning proves to be a powerful approach for modelling complex
environmental data, offering accurate, data-driven insights that can support timely and
informed decision-making for public health and urban planning.

5.2 Future Work

There are a number of exciting avenues for future machine learning-based air quality
prediction research. One crucial area is the incorporation of indoor air quality data,
which is becoming more and more crucial for a more thorough evaluation of human
exposure. A comprehensive understanding of the effects of pollution can be obtained by
combining measurements of indoor and outdoor air quality. Furthermore, the spatial and
temporal resolution of predictions can be greatly enhanced by combining data on traffic
flow, industrial activity records, urban infrastructure, and real-time weather forecasts.

57
The model's capacity to identify intricate patterns and correlations in the dynamics of
air quality can be improved by utilizing sophisticated deep learning architectures such
attention mechanisms, transformer-based models, and graph neural networks.
Furthermore, using IoT-based sensor networks and satellite imaging will allow for
extensive, real-time air quality monitoring. Building trust and giving policymakers and
urban planners useful information will also require the development of models that
incorporate explainability and uncertainty estimation. Lastly, developing transferable
models that are applicable to other climates, seasons, and geographical areas is still a
crucial objective for developing reliable and scalable air quality forecast systems.

58
References
[1] Z. Zhao, J. Wu, F. Cai, S. Zhang, and Y. G. Wang, “A statistical learning
framework for spatial-temporal feature selection and application to air quality
index forecasting,” Ecol Indic, vol. 144, Nov. 2022, doi:
10.1016/[Link].2022.109416.
[2] N. N. Maltare and S. Vahora, “Air Quality Index prediction using machine
learning for Ahmedabad city,” Digital Chemical Engineering, vol. 7, Jun. 2023,
doi: 10.1016/[Link].2023.100093.
[3] S. K. Natarajan, P. Shanmurthy, D. Arockiam, B. Balusamy, and S. Selvarajan,
“Optimized machine learning model for air quality index prediction in major
cities in India,” Sci Rep, vol. 14, no. 1, Dec. 2024, doi: 10.1038/s41598-024-
54807-1.
[4] A. Binbusayyis, M. A. Khan, M. M. Ahmed A, and W. R. S. Emmanuel, “A
deep learning approach for prediction of air quality index in smart city,”
Discover Sustainability, vol. 5, no. 1, Dec. 2024, doi: 10.1007/s43621-024-
00272-9.
[5] J. Wang, X. Li, L. Jin, J. Li, Q. Sun, and H. Wang, “An air quality index
prediction model based on CNN-ILSTM,” Sci Rep, vol. 12, no. 1, Dec. 2022,
doi: 10.1038/s41598-022-12355-6.
[6] R. Zayed and M. Abbod, “Breathable Cities: Dynamic Machine Learning
Modelling Approaches for Advanced Air Pollution Control,” Applied Sciences
(Switzerland), vol. 14, no. 13, Jul. 2024, doi: 10.3390/app14135581.
[7] Q. Liu, B. Cui, and Z. Liu, “Air Quality Class Prediction Using Machine
Learning Methods Based on Monitoring Data and Secondary Modeling,”
Atmosphere (Basel), vol. 15, no. 5, May 2024, doi: 10.3390/atmos15050553.
[8] K. Karthick, S. K. Aruna, R. Dharmaprakash, and G. Ravindiran, “Integrating
machine learning techniques for Air Quality Index forecasting and insights from
pollutant-meteorological dynamics in sustainable urban environments,” Earth
Sci Inform, Aug. 2024, doi: 10.1007/s12145-024-01382-8.
[9] M. Emeç and M. Yurtsever, “A novel ensemble machine learning method for
accurate air quality prediction,” International Journal of Environmental Science
and Technology, 2024, doi: 10.1007/s13762-024-05671-z.
[10] L. Chen et al., “Deep Citywide Multisource Data Fusion-Based Air Quality
Estimation,” IEEE Trans Cybern, vol. 54, no. 1, pp. 111–122, Jan. 2024, doi:
10.1109/TCYB.2023.3245618.
[11] Y. Cao, D. Zhang, S. Ding, W. Zhong, and C. Yan, “A Hybrid Air Quality
Prediction Model Based on Empirical Mode Decomposition,” 2024.
[12] Z. Sadriddin, R. R. Mekuria, and M. S. Gaso, “Machine Learning Models for
Advanced Air Quality Prediction,” in ACM International Conference

59
Proceeding Series, Association for Computing Machinery, Jun. 2024, pp. 51–
56. doi: 10.1145/3674912.3674915.
[13] K. Chatterjee et al., “Toward Cleaner Industries: Smart Cities’ Impact on
Predictive Air Quality Management,” IEEE Access, vol. 12, pp. 78895–78910,
2024, doi: 10.1109/ACCESS.2024.3406502.
[14] A. K. Rad, S. O. Razmi, M. J. Nematollahi, A. Naghipour, F. Golkar, and M.
Mahmoudi, “Machine learning models for predicting interactions between air
pollutants in Tehran Megacity, Iran,” Alexandria Engineering Journal, vol. 104,
pp. 464–479, Oct. 2024, doi: 10.1016/[Link].2024.08.023.
[15] M. Elsarraj, Y. Mahmoudi, and A. Keshmiri, “Quantifying indoor infection risk
based on a metric-driven approach and machine learning,” Build Environ, vol.
251, Mar. 2024, doi: 10.1016/[Link].2024.111225.
[16] M. A. Alolayan, A. Almutairi, S. M. Aladwani, and S. Alkhamees,
“Investigating major sources of air pollution and improving spatiotemporal
forecast accuracy using supervised machine learning and a proxy,” Journal of
Engineering Research (Kuwait), vol. 11, no. 3, pp. 87–93, Sep. 2023, doi:
10.1016/[Link].2023.100126.
[17] Q. Shao, J. Chen, and T. Jiang, “A Novel Coupled Optimization Prediction
Model for Air Quality,” IEEE Access, vol. 11, pp. 69667–69685, 2023, doi:
10.1109/ACCESS.2023.3293249.
[18] S. Berkani, I. Gryech, M. Ghogho, B. Guermah, and A. Kobbane, “Data Driven
Forecasting Models for Urban Air Pollution: MoreAir Case Study,” IEEE
Access, vol. 11, pp. 133131–133142, 2023, doi:
10.1109/ACCESS.2023.3331565.
[19] D. Iskandaryan, F. Ramos, and S. Trilles, “Graph Neural Network for Air
Quality Prediction: A Case Study in Madrid,” IEEE Access, vol. 11, pp. 2729–
2742, 2023, doi: 10.1109/ACCESS.2023.3234214.
[20] S. J. Livingston, S. D. Kanmani, A. S. Ebenezer, D. Sam, and A. Joshi, “An
ensembled method for air quality monitoring and control using machine
learning,” Measurement: Sensors, vol. 30, Dec. 2023, doi:
10.1016/[Link].2023.100914.
[21] F. Farhadi, R. Palacin, and P. Blythe, “Machine Learning for Transport Policy
Interventions on Air Quality,” IEEE Access, vol. 11, pp. 43759–43777, 2023,
doi: 10.1109/ACCESS.2023.3272662.
[22] S. Al-Eidi, F. Amsaad, O. Darwish, Y. Tashtoush, A. Alqahtani, and N.
Niveshitha, “Comparative Analysis Study for Air Quality Prediction in Smart
Cities Using Regression Techniques,” IEEE Access, vol. 11, pp. 115140–
115149, 2023, doi: 10.1109/ACCESS.2023.3323447.
[23] B. Alam, A. Hussain, and M. Fayaz, “An Effective Approach for Air Quality
Prediction in Bishkek Based on Machine Learning techniques,” in ACM

60
International Conference Proceeding Series, Association for Computing
Machinery, Oct. 2023, pp. 42–47. doi: 10.1145/3633598.3633606.
[24] J. Han, H. Liu, H. Xiong, and J. Yang, “Semi-Supervised Air Quality
Forecasting via Self-Supervised Hierarchical Graph Neural Network,” IEEE
Trans Knowl Data Eng, vol. 35, no. 5, pp. 5230–5243, May 2023, doi:
10.1109/TKDE.2022.3149815.
[25] N. H. Motlagh et al., “Unmanned Aerial Vehicles for Air Pollution Monitoring:
A Survey,” IEEE Internet Things J, vol. 10, no. 24, pp. 21687–21704, Dec.
2023, doi: 10.1109/JIOT.2023.3290508.
[26] I. N. K. Wardana, S. A. Fahmy, and J. W. Gardner, “TinyML Models for a Low-
Cost Air Quality Monitoring Device,” IEEE Sens Lett, vol. 7, no. 11, Nov. 2023,
doi: 10.1109/LSENS.2023.3315249.
[27] J. Kalajdjieski, K. Trivodaliev, G. Mirceva, S. Kalajdziski, and S. Gievska, “A
Complete Air Pollution Monitoring and Prediction Framework,” IEEE Access,
vol. 11, pp. 88730–88744, 2023, doi: 10.1109/ACCESS.2023.3251346.
[28] J. Dong, Y. Zhang, and J. Hu, “Short-term air quality prediction based on EMD-
transformer-BiLSTM,” Sci Rep, vol. 14, no. 1, Dec. 2024, doi: 10.1038/s41598-
024-67626-1.
[29] X. B. Jin et al., “Deep Spatio-Temporal Graph Network with Self-Optimization
for Air Quality Prediction,” Entropy, vol. 25, no. 2, Feb. 2023, doi:
10.3390/e25020247.
[30] J. Luo and Y. Gong, “Air pollutant prediction based on ARIMA-WOA-LSTM
model,” Atmos Pollut Res, vol. 14, no. 6, Jun. 2023, doi:
10.1016/[Link].2023.101761.
[31] K. Zhang, X. Yang, H. Cao, J. Thé, Z. Tan, and H. Yu, “Multi-step forecast of
PM2.5 and PM10 concentrations using convolutional neural network integrated
with spatial–temporal attention and residual learning,” Environ Int, vol. 171,
Jan. 2023, doi: 10.1016/[Link].2022.107691.
[32] G. Ravindiran, G. Hayder, K. Kanagarathinam, A. Alagumalai, and C. Sonne,
“Air quality prediction by machine learning models: A predictive study on the
indian coastal city of Visakhapatnam,” Chemosphere, vol. 338, Oct. 2023, doi:
10.1016/[Link].2023.139518.
[33] L. Mampitiya et al., “Machine Learning Techniques to Predict the Air Quality
Using Meteorological Data in Two Urban Areas in Sri Lanka,” Environments -
MDPI, vol. 10, no. 8, Aug. 2023, doi: 10.3390/environments10080141.
[34] S. Kumari and S. K. Singh, “Machine learning-based time series models for
effective CO2 emission prediction in India,” Environmental Science and
Pollution Research, vol. 30, no. 55, pp. 116601–116616, Nov. 2023, doi:
10.1007/s11356-022-21723-8.

61
[35] G. I. Drewil and R. J. Al-Bahadili, “Air pollution prediction using LSTM deep
learning and metaheuristics algorithms,” Measurement: Sensors, vol. 24, Dec.
2022, doi: 10.1016/[Link].2022.100546.
[36] A. Gilik, A. S. Ogrenci, and A. Ozmen, “Air quality prediction using
CNN+LSTM-based hybrid deep learning architecture,” Environmental Science
and Pollution Research, vol. 29, no. 8, pp. 11920–11938, Feb. 2022, doi:
10.1007/s11356-021-16227-w.
[37] A. Heydari, M. Majidi Nezhad, D. Astiaso Garcia, F. Keynia, and L. De Santoli,
“Air pollution forecasting application based on deep learning model and
optimization algorithm,” Clean Technol Environ Policy, vol. 24, no. 2, pp. 607–
621, Mar. 2022, doi: 10.1007/s10098-021-02080-5.
[38] B. Y. Kim, Y. K. Lim, and J. W. Cha, “Short-term prediction of particulate
matter (PM10 and PM2.5) in Seoul, South Korea using tree-based machine
learning algorithms,” Atmos Pollut Res, vol. 13, no. 10, Oct. 2022, doi:
10.1016/[Link].2022.101547.
[39] E. Gladkova and L. Saychenko, “Applying machine learning techniques in air
quality prediction,” in Transportation Research Procedia, Elsevier B.V., 2022,
pp. 1999–2006. doi: 10.1016/[Link].2022.06.222.
[40] X. Lin, H. Wang, J. Guo, and G. Mei, “A Deep Learning Approach Using Graph
Neural Networks for Anomaly Detection in Air Quality Data Considering
Spatiotemporal Correlations,” IEEE Access, vol. 10, pp. 94074–94088, 2022,
doi: 10.1109/ACCESS.2022.3204284.
[41] X. Yi, Z. Duan, R. Li, J. Zhang, T. Li, and Y. Zheng, “Predicting Fine-Grained
Air Quality Based on Deep Neural Networks,” IEEE Trans Big Data, vol. 8, no.
5, pp. 1326–1339, Oct. 2022, doi: 10.1109/TBDATA.2020.3047078.
[42] T. Baldi, G. Delnevo, R. Girau, and S. Mirri, “On the Prediction of Air Quality
within Vehicles using Outdoor Air Pollution: Sensors and Machine Learning
Algorithms,” in NET4us 2022 - Proceedings of the ACM SIGCOMM Workshop
on Networked Sensing Systems for Sustainable Society, Association for
Computing Machinery, Inc, Aug. 2022, pp. 14–19. doi:
10.1145/3538393.3544934.
[43] J. Wang, L. Jin, X. Li, S. He, M. Huang, and H. Wang, “A Hybrid Air Quality
Index Prediction Model Based on CNN and Attention Gate Unit,” IEEE Access,
vol. 10, pp. 113343–113354, 2022, doi: 10.1109/ACCESS.2022.3217242.
[44] D. Fister, J. Pérez-Aracil, C. Peláez-Rodríguez, J. Del Ser, and S. Salcedo-Sanz,
“Accurate long-term air temperature prediction with Machine Learning models
and data reduction techniques,” Appl Soft Comput, vol. 136, Mar. 2023, doi:
10.1016/[Link].2023.110118.
[45] S. Hameed, A. Islam, K. Ahmad, S. B. Belhaouari, J. Qadir, and A. Al-Fuqaha,
“Deep learning based multimodal urban air quality prediction and traffic

62
analytics,” Sci Rep, vol. 13, no. 1, Dec. 2023, doi: 10.1038/s41598-023-49296-
7.
[46] M. S. Ramadan, A. Abuelgasim, A. H. Almurshidi, and N. Al Hosani, “A
comprehensive spatiotemporal approach to mapping air quality distribution and
prediction in desert region,” Urban Clim, vol. 58, Nov. 2024, doi:
10.1016/[Link].2024.102137.
[47] B. Liu, Z. Qi, and L. Gao, “Enhanced Air Quality Prediction through Spatio-
temporal Feature Sxtraction and Fusion: A Self-tuning Hybrid Approach with
GCN and GRU,” Water Air Soil Pollut, vol. 235, no. 8, Aug. 2024, doi:
10.1007/s11270-024-07346-4.
[48] Z. Liu, D. Ji, and L. Wang, “PM2.5 concentration prediction based on EEMD-
ALSTM,” Sci Rep, vol. 14, no. 1, Dec. 2024, doi: 10.1038/s41598-024-63620-
9.
[49] J. Duan, Y. Gong, J. Luo, and Z. Zhao, “Air-quality prediction based on the
ARIMA-CNN-LSTM combination model optimized by dung beetle optimizer,”
Sci Rep, vol. 13, no. 1, Dec. 2023, doi: 10.1038/s41598-023-36620-4.
[50] X. Zhang, X. Jiang, and Y. Li, “Prediction of air quality index based on the
SSA-BiLSTM-LightGBM model,” Sci Rep, vol. 13, no. 1, Dec. 2023, doi:
10.1038/s41598-023-32775-2.
[51] S. Chennareddy, S. Saha, A. Das, and T. Kayal, “PM2.5 Concentration
Forecasting in the Kolkata Region With Spatiotemporal Sliding Window
Approaches,” IEEE Access, vol. 12, pp. 82333–82353, 2024, doi:
10.1109/ACCESS.2024.3411777.
[52] P. Mottahedin, B. Chahkandi, R. Moezzi, A. M. Fathollahi-Fard, M. Ghandali,
and M. Gheibi, “Air Quality Prediction and Control Systems Using Machine
Learning and Adaptive Neuro-Fuzzy Inference System,” Heliyon, p. e39783,
Oct. 2024, doi: 10.1016/[Link].2024.e39783.
[53] A. T. Nguyen, D. H. Pham, B. L. Oo, Y. Ahn, and B. T. H. Lim, “Predicting air
quality index using attention hybrid deep learning and quantum-inspired
particle swarm optimization,” J Big Data, vol. 11, no. 1, Dec. 2024, doi:
10.1186/s40537-024-00926-5.
[54] Z. Zhao, J. Wu, F. Cai, S. Zhang, and Y. G. Wang, “A hybrid deep learning
framework for air quality prediction with spatial autocorrelation during the
COVID-19 pandemic,” Sci Rep, vol. 13, no. 1, Dec. 2023, doi: 10.1038/s41598-
023-28287-8.
[55] P. Dey, S. Dev, and B. S. Phelan, “Predicting Multivariate Air Pollution: A
Gaussian-Mixture Nested Factorial Variational Autoencoder Approach,” IEEE
Geoscience and Remote Sensing Letters, vol. 21, 2024, doi:
10.1109/LGRS.2024.3416343.

63
[56] C. Liu, G. Pan, D. Song, and H. Wei, “Air Quality Index Forecasting via
Genetic Algorithm-Based Improved Extreme Learning Machine,” IEEE Access,
vol. 11, pp. 67086–67097, 2023, doi: 10.1109/ACCESS.2023.3291146.
[57] T. Emmanuel, T. Maupong, D. Mpoeleng, T. Semong, B. Mphago, and O.
Tabona, “A survey on missing data in machine learning,” J Big Data, vol. 8, no.
1, Dec. 2021, doi: 10.1186/s40537-021-00516-9.
[58] D. J. Zamanzadeh et al., “Autopopulus: A Novel Framework for Autoencoder
Imputation on Large Clinical Datasets on behalf of the CURE-CKD Study
team.” [Online]. Available: [Link]
[59] S. Lee, H. Park, and H.-C. Kim, “Fraud Detection on Crowdfunding Platforms
Using Multiple Feature Selection Methods,” IEEE Access, vol. 13, pp. 40133–
40148, 2025, doi: 10.1109/ACCESS.2025.3547396.
[60] Y. Dai et al., “Retrieval of Land Surface Temperature From Passive Microwave
Observations Using CatBoost-Based Adaptive Feature Selection,” IEEE J Sel
Top Appl Earth Obs Remote Sens, 2025, doi: 10.1109/JSTARS.2025.3532605.

64

Common questions

Powered by AI

CatBoost-GRU maintains cost-effectiveness and scalability in AQI forecasting by achieving high predictive accuracy with relatively low computational and training resource demands. It requires less training time and incurs lower computational costs compared to models like CatBoost-TCN and CNN-GRU, due to its efficient architecture and optimized feature handling. For example, its training time is substantially reduced, making it suitable for real-time applications where frequent retraining is necessary. The model's ability to deliver accurate predictions without heavy computational load makes it a practical and scalable solution for large-scale deployments .

CatBoost-GRU is a better choice for real-time air quality forecasting scenarios as it offers significantly faster training times and lower computational costs compared to CNN-GRU. CatBoost-GRU requires only 365.2 seconds to train on the Autoencoder dataset versus CNN-GRU's 3314.83 seconds, and it incurs much less computational expense (31,057 units vs 88,001 units). This efficiency makes CatBoost-GRU highly suitable for environments requiring rapid retraining and deployment, effectively supporting frequent data updates and scalability in real-time contexts .

The performance comparison of CatBoost-GRU and LSTM across various datasets suggests that integrating boosting techniques with RNNs can significantly enhance AQI prediction. CatBoost-GRU consistently outperforms LSTM, with lower RMSE and higher R2 scores, demonstrating better predictive accuracy and fit to data. For example, on the Autoencoder dataset, CatBoost-GRU has an RMSE of 20.9267 compared to LSTM's 36.3154, and an R2 of 0.9835 against 0.9502. This shows that the combination of gradient boosting with recurrent neural networks leverages structured features more effectively, leading to improved handling of temporal dynamics and complex data patterns essential for AQI forecasting .

Choosing between CatBoost-GRU and CNN-GRU for AQI prediction involves considering trade-offs between predictive accuracy and computational efficiency. CatBoost-GRU offers slightly better prediction accuracy with a lower RMSE and higher R2 score. For instance, on the KNN dataset, it has an RMSE of 24.1831 and an R2 of 0.978. However, CNN-GRU, while slightly behind in accuracy, still delivers reasonable performance with a narrower MAE difference on the KNN dataset. Importantly, CatBoost-GRU's significantly lower training time and computational cost (31,057 units vs 88,001 units) may make it preferable in scenarios where resource efficiency and rapid model updates are needed. The decision will depend on the specific needs of the application, such as the importance of speed and cost versus marginal improvements in accuracy .

The integration of CatBoost with GRU contributes to enhanced predictive accuracy in air quality forecasting by combining the strengths of both approaches: CatBoost’s powerful decision tree-based ensemble learning with GRU’s ability to handle sequential data and temporal dependencies. CatBoost efficiently captures nonlinear relationships and interactions among features, while GRU learns temporal patterns effectively. This synergy results in more accurate and reliable predictions, as evidenced by CatBoost-GRU's lower RMSE and higher R2 scores across datasets compared to traditional RNN models like LSTM. The boosted structure of CatBoost also aids in robust generalization across diverse AQI levels, making it well-suited for complex forecasting tasks .

The CatBoost-GRU hybrid model demonstrates superior performance over the CNN-LSTM model across both the Autoencoder and KNN datasets. On the Autoencoder dataset, CatBoost-GRU achieves a significantly lower RMSE of 20.9267 compared to CNN-LSTM's 35.8272, indicating a more accurate prediction ability. The MSE and MAE metrics also reinforce this, with CatBoost-GRU having values of 437.9258 and 11.3266, respectively, much lower than those of CNN-LSTM. Additionally, CatBoost-GRU's training time is shorter and computational cost lower, making it more efficient for practical implementations .

While the CatBoost-TCN model may offer a slightly more accurate prediction under certain conditions, its computational overhead is significantly higher, making it less efficient compared to CatBoost-GRU. The CatBoost-TCN incurs a computational expense of 93,889 units versus CatBoost-GRU's 31,633 units, with limited gains in predictive accuracy. This results in a scenario where the CatBoost-GRU model is favored for its balance of accuracy, versatility, and cost-effectiveness, particularly in resource-constrained environments .

Feature selection techniques play a crucial role in enhancing the performance of the CatBoost-GRU model for AQI prediction by improving model efficiency and accuracy. Using the autoencoder-based feature selection, CatBoost-GRU attains lower RMSE and MSE values compared to when KNN-based features are used, indicating more accurate and precise predictions. The autoencoder-based feature selection leads to a more compact and effective representation of the dataset, which enhances the model's ability to learn relevant patterns and perform robustly across different AQI ranges. This impact is reflected in the higher R2 scores achieved by the model on the autoencoder dataset, demonstrating better capture of data variability and underlying structures .

The CatBoost-GRU model significantly outperforms the ILSTM model in terms of prediction accuracy for AQI forecasting using the Autoencoder dataset. It achieves a lower RMSE of 20.9267 compared to ILSTM's 30.3574, which indicates more precise predictions. Moreover, CatBoost-GRU's R2 score of 0.9835 is higher than ILSTM's 0.9652, demonstrating a better fit to the dataset and more effective capture of data patterns .

On the KNN dataset, CatBoost-GRU demonstrates its effectiveness over LSTM by achieving a lower RMSE of 24.1831 compared to LSTM's 34.5155, which indicates better prediction accuracy even in severe AQI cases. The model's ability to better capture data patterns is further evidenced by its higher R2 score of 0.978 compared to LSTM's 0.9551. This suggests that CatBoost-GRU can recognize extreme AQI levels more accurately, providing a more robust solution for scenarios where timely and precise detection of high pollution levels is crucial .

You might also like