Main
Main
Keywords: Ambient air pollution is a pervasive global issue that poses significant health risks. Among pollutants, ozone
Pollution (O3 ) is responsible for an estimated 1 to 1.2 million premature deaths yearly. Furthermore, O3 adversely affects
Time series forecasting climate warming, crop productivity, and more. Its formation occurs when nitrogen oxides and volatile organic
Ozone
compounds react with short-wavelength solar radiation. Consequently, urban areas with high traffic volume
Feature selection
and elevated temperatures are particularly prone to elevated O3 levels, which pose a significant health risk to
Deep learning
XAI
their inhabitants. In response to this problem, many countries have developed web and mobile applications that
provide real-time air pollution information using sensor data. However, while these applications offer valuable
insight into current pollution levels, predicting future pollutant behavior is crucial for effective planning and
mitigation strategies. Therefore, our main objectives are to develop accurate and efficient prediction models
and identify the key factors that influence O3 levels. We adopt a time series forecasting approach to address
these objectives, which allows us to analyze and predict future O3 behavior. Additionally, we tackle the feature
selection problem to identify the most relevant features and periods that contribute to prediction accuracy by
introducing a novel method called the Time Selection Layer in Deep Learning models, which significantly
improves model performance, reduces complexity, and enhances interpretability. Our study focuses on data
collected from five representative areas in Seville, Cordova, and Jaen provinces in Spain, using multiple
sensors to capture comprehensive pollution data. We compare the performance of three models: Lasso, Decision
Tree, and Deep Learning with and without incorporating the Time Selection Layer. Our results demonstrate
that including the Time Selection Layer significantly enhances the effectiveness and interpretability of Deep
Learning models, achieving an average effectiveness improvement of 9% across all monitored areas.
∗ Corresponding author.
E-mail addresses: mjimenez3@[Link] (M.J. Jiménez-Navarro), mariamartinez@[Link] (M. Martínez-Ballesteros), fmaralv@[Link] (F. Martínez-Álvarez),
guaasecor@[Link] (G. Asencio-Cortés).
[Link]
Received 31 July 2023; Received in revised form 1 February 2024; Accepted 17 February 2024
Available online 21 March 2024
1568-4946/© 2024 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license ([Link]
nc-nd/4.0/).
M.J. Jiménez-Navarro et al. Applied Soft Computing 157 (2024) 111504
accurate information to the population of affected areas and alert them ones. Although this approach is often efficient, it does not guarantee
when there are exceptionally high levels [12]. In addition, understand- the selection of the most relevant features. It may not necessarily cor-
ing the behavior of the pollutants and the influent variables has been respond to the importance within the assigned model, which limits the
extensively studied [13–15]. reliability in interpreting how the model uses these selected features.
Several models have been proposed for time series forecasting [16], The wrapper approach employs the model itself as a heuristic,
achieving remarkable results in various problems. However, there is a retraining it with different feature sets. This approach can yield optimal
growing demand for accurate predictions and computationally efficient feature selections for a specific model depending on the process and
and interpretable solutions. Computation resources can be reduced may enhance interpretability by aligning the features more closely to
using hardware- and/or software-centric techniques [17], promoting the model. However, it comes with the drawback of potentially being
the development of green AI [18]. Various approaches exist to improve computationally intensive, making it challenging to scale for a larger
model interpretability in artificial intelligence, such as global/local number of features and complex models.
model-agnostic or example-based methods. In this context, feature The embedded approach usually offers the best approach, com-
selection [19] is a unique technique that offers benefits in reducing bining the strengths of the filter and wrapper methods by efficiently
computational costs and improving interpretability. This technique selecting features that are both optimal for the model and interpretable.
aims to identify a subset of relevant features that are sufficient or more Furthermore, this feature selection approach usually does not require
informative than the original set of features. This technique reduces hyperparameter tuning in contrast to other approaches, simplifying the
computational resources regarding memory and execution time, and training process and reducing computational costs. However, imple-
the solution becomes better understood by focusing on the relevant
menting these methods involves developing an algorithm that incor-
identified features. Furthermore, in time series forecasting problems,
porates the feature selection procedure into its design.
providing information about the relevant features along the temporal
dimension is essential.
In this work, we study a dataset collected from six areas of Seville 2.1. Filter
and Jaen provinces, Spain. Our study focuses on models incorporating
a global interpretability method, which embeds the feature selection In their study, Khemphila et al. [26] utilize information gain scoring
mechanism within the model design instead of other feature selection to rank features in a heart disease dataset. A threshold-based feature
approaches. Thus, we can determine future behavior and relevant fea- selection is applied, and a neural network is trained with the selected
tures once the model is obtained. Examples of models that incorporate data. Unlike their approach, our work focuses on embedded methods,
feature selection include tree-based methods [20] (e.g., Decision Tree wherein the selected features are directly related to the trained model.
or CART) and regularization-based methods [21] (e.g., Lasso or L1). Lui et al. [27] construct an ensemble of neural networks grouped
Neural networks may be a suitable approach for time series fore- into threes. Each group employs different selection methods, includ-
casting [22]. However, two main issues remain open: computational ing PCA, RankSum, and t-test. The group outputs are averaged, and
cost and interpretability. The Temporal Selection Layer (TSL) proposed majority voting is used for final predictions. This method significantly
in Jimenez-Navarro et al. [23,24] addresses both of these issues by impacts efficacy but performs poorly due to indiscriminate feature
transforming a neural network into a model with embedded feature selection.
selection which attempts to improve the efficacy and interpretability Yan [28] adopts an AdaBoost model to identify relevant features
while reducing computational costs1 . and trains a neural network using these features. While effective for
Therefore, this study presents the following contributions: stock index prediction, this method shares a typical challenge with
filter methods: the feature selection process remains separate from the
• An improvement of the originally proposed TSL by adapting model, potentially leading to suboptimal feature choices for the neural
regularization for multi-step forecasting problems.
network.
• Conducted a real-world scenario comparison of Decision Tree
Niu et al. [29] employ the non-linear feature selection algorithm
(DT), L1, Feed-Forward (FF), and Long Short-Term Memory
RReliefF. The relevant features this algorithm determines are then used
(LSTM) neural networks. The selected models embed a feature se-
as input for neural networks, resulting in considerable performance
lection process and are known for their simplicity and efficiency.
improvements across all scenarios.
This study is outlined as follows. Relevant and previous work are
reviewed in Section 2. In Section 3, the proposed improvement is 2.2. Wrapper
introduced along with the experimental workflow. The main results of
the performance of Section 4 and their comparison with the rest of Verikas et al. [30] introduce a feature ranking method based on
the algorithms are discussed. Finally, Section 5 summarizes the main cross-entropy error constraints. They recursively remove features from
conclusions drawn from this study. least to most important during neural network training. Our paper fo-
cuses on efficacy and performance, aiming to minimize computational
2. Related work resources during feature selection.
Kabir et al. [31] propose an incremental feature selection and
Air pollution is a well-known problem with several proposed solu- model-building method. They use feature correlations to divide features
tions, but not all include a feature selection process. In this section, into groups and incrementally add them during training if they improve
we analyze several approaches to feature selection for neural net- error rates.
works to justify their effectiveness in a wide range of problems. The Tong et al. [32] use a genetic algorithm in which individuals
approaches have been categorized into three types [25] of feature represent models, leading to high memory consumption for numerous
selection approaches: filter, wrapper, and embedded.
populations. The computational cost can become prohibitive with large
The filter approach relies on a heuristic to assess the importance
models or datasets, making it unsuitable for our paper’s scope.
of each feature and then establishes a threshold to remove irrelevant
Alshammari et al. [33] utilize a neural network and the Mayfly
metaheuristic algorithm to optimize the number of relevant features.
1
All the code needed to reproduce the results has been included in the fol- In contrast, our method proposes an efficient and effective embedded
lowing repository: [Link] along with feature selection approach, avoiding the need for a significant increase
detailed instructions. in computational resources.
2
M.J. Jiménez-Navarro et al. Applied Soft Computing 157 (2024) 111504
Fig. 1. Summary of experimentation workflow. The data for each station 𝑆𝑖 is transformed, used for the training, and finally evaluated. During the preprocessing, the data is
𝑡𝑒𝑠𝑡 𝑡𝑟𝑎𝑖𝑛
divided in different folds returning the test 𝑋𝑖,𝑦 , corresponding to the year 𝑦, and training dataset 𝑋𝑖,𝑦 with the remaining years for station 𝑖. Then, the missing data is inputted,
the data is standardized using the 𝑋𝑖𝑡𝑟𝑎𝑖𝑛 , and the lagged information is built. Furthermore, the training process is repeated for each preprocessed data 𝑋𝑖,𝑦 𝑡𝑟𝑎𝑖𝑛
, selected model 𝑀𝑗 ,
and sampled hyperparameters 𝐻𝑘𝑗 from the Bayesian optimization for 𝐾 iterations. Finally, the model is evaluated for the 𝑋𝑖,𝑦
𝑡𝑒𝑠𝑡
to obtain the results.
2.3. Embedded Firstly, we propose two simple and very efficient models that also
provide an embedded feature selection. Decision tree (DT) and Lasso
Cancela et al. [34] introduce the E2E-FS technique for integrating (L1) models were chosen due to their well-known simplicity and high
feature selection within neural networks. This method employs an efficiency. In this case, both models use a direct approach for multi-step
additional loss function strategy to effectively filter out irrelevant fea- time series forecasting.
tures. In contrast, our method automatically determines the number of Feed-Forward neural networks (FF) and Long Short-Term Memory
features through an approximation of the Heaviside function to ensure (LSTM) [38] were used by employing the direct approach to produce
differentiability without modifying the loss function. multi-step forecasting. Note that neural networks are not usually con-
Da Costa et al. [35] propose a feature selection method that en- sidered simple in other problems. However, in our experience, neural
hances the stability of feature rankings applied to neural networks. networks do not need to be particularly deep and wide to solve time
However, our focus is automatically selecting the best subset of features series forecasting problems [39,40]. Therefore, we decided to employ
to improve efficacy or achieve comparable results with fewer features. this model type as it accomplishes the specified criterion in this context.
Our approach does not rely on a two-step process involving input layer
weight analysis and subsequent feature ranking. 3.2. Time selection layer
Yuan et al. [36] present an embedded feature selection method
for neural networks using the point-centered convolutional neural net- The FF and LSTM neural networks are employed in two ways:
work. While effective in identifying moldy peanuts using hyperspectral
including and not including the TSL. For instance, when the TSL is
images, their approach assigns continuous weights, making determin-
used, the model is named TLSTM in the case of LSTM and TFF for
ing the relevance of specific features challenging.
FF. This layer is formally defined as the element-wise multiplication
Zhang et al. [37] propose incorporating feature selection within a
of the binarized weights and the input tensor, defined in the following
neural network framework using the Group Lasso penalty. Although
equation:
effective for Feed-Forward layers, this approach limits its use in other
models. On the contrary, our research offers a versatile and general- 𝑇 𝑆𝐿(𝑋 𝑀𝑥𝐷 ) = 𝐻(𝑊 𝑀𝑥𝐷 )◦𝑋 𝑀𝑥𝐷 (1)
purpose approach that applies to various neural network architectures.
where ‘‘◦’’ represents the Hadamard product, 𝑊 represents the window
size of the time series, and 𝐷 is the dimension size. Moreover, an
3. Methodology
approximation to the Heaviside function (𝐻) serves as a gate that
indicates if the feature is selected.
This section details the methodology proposed to evaluate the dif-
Using backpropagation, TSL is trained as an additional layer at the
ferent approaches. Section 3.1 shows the proposed models employed
top of any neural network. When a weight in the feature selection mask
in the comparison. Section 3.2 describes TSL and its background. Sec-
is set to zero, it eliminates the contribution of the corresponding feature
tion 3.3 shows the improvement applied to TSL in this work. Section 3.4
in the forward step.
shows the data used during the experimentation and its properties.
Weight regularization is essential in this layer, applied as a Lasso re-
Section 3.5 provides an overview of how the experimentation workflow
gression to the feature selection mask to prevent undesirable outcomes
is performed.
such as not selecting any feature. This regularization is calculated as
Fig. 1 summarizes the workflow followed during the experimenta-
the sum of a regularization factor multiplied by the number of selected
tion to obtain the results.
features. This regularization is then included in the loss of the model.
3.1. Models
3.3. Regularization improvement in time selection layer
For comparative purposes, we established the following criteria for
selecting the models: We present an extension of our previous work [23] with an ad-
ditional improvement. Our proposal introduces TSL, an extra layer
• Models should have efficient inference time. incorporating embedded feature selection into a neural network. TSL
• Models should be interpretable, meaning they should indicate layer was previously used with remarkable results in time series fore-
relevant features. casting. However, it was not utilized for multi-step forecasting. We
• Simplicity is preferred, as discussed in Section 4, where we ob- hypothesize that forecasting larger periods requires more lagged infor-
served that simpler models yield remarkable results for this prob- mation to reproduce future behavior effectively. Therefore, the need for
lem. additional lagged features becomes apparent as the horizon increases.
3
M.J. Jiménez-Navarro et al. Applied Soft Computing 157 (2024) 111504
Table 1
Table with datasets relevant properties employed in all areas.
Property Aljarafe Asomadilla Bermejales Torneo Ronda del Valle
Location Seville Cordova Seville Seville Jaen
Area type Suburban Suburban Urban Urban Urban
Emission source Background Background Background Traffic Background
Features TMP, WD, WP, CO, NO2 , 𝑃 𝑀10 and O3
Sample frequency Hourly
Sample period 2006–2023
Ozone unit μg∕m3
3.4. Data
Hourly O3 multivariate time series were obtained annually from During cross-validation, each data split is standardized by trans-
2006 to 2023. Data were collected from four monitoring stations in the forming the data to have a zero mean and a standard deviation of one.
metropolitan area of Seville and one in Jaen. These monitoring sites The mean and standard deviation of the training set are calculated and
were classified according to the type of area (U-Urban, S-Suburban) and used to standardize the validation and test sets.
the predominant emission source (B-Background, T-Traffic). The sites Moreover, when dealing with time series problems, it is necessary to
analyzed were Aljarafe (S-B), Bermejales (U-B), Torneo (U-T) in Seville, incorporate lagged temporal information as input for the models. This
Asomadilla (S-B) in Cordova, and Ronda del Valle (U-B) in Jaen. This means transforming the original time series 𝑋 𝑁,𝐷 with 𝑁 samples and
selection allowed for considering various meteorological conditions and 𝐷 features into 𝑋 𝑁,𝐷,𝑊 , where 𝑊 represents the number of lagged in-
local factors that influence the behavior of O3 . Meteorological condi- formation (window size) included for each sample and feature. Finally,
tions included average temperature (TMP), wind direction (WD), and since the goal is to predict the future, the matrix 𝑌 𝑁,𝐻 is constructed
wind speed (WS). The local factors considered were carbon monoxide for each sample, providing 𝐻 future Ozone (O3 ) values that constitute
(CO), nitrogen dioxide (NO2 ), and particles of 10 μm or less (𝑃 𝑀10 ). the prediction.
The O3 measurements were obtained using the reference monitoring
method [41].
3.5.2. Training
Table 1 summarizes the properties of the datasets employed. Each selected model has its own set of hyperparameters that greatly
influence its effectiveness. A search algorithm is used to determine the
3.5. Experimentation process best set of hyperparameters, which aims to find the optimal configura-
tion across all the tested folds.
Due to the differences in the behavior of each zone, we decided to The chosen search algorithm is Bayesian search [43], which has
model each one individually. The same methodology was applied in shown outstanding performance in recent studies [44,45]. In each
each zone, divided into preprocessing, training, and evaluation. iteration, the algorithm generates a set of hyperparameters. The model
is then parameterized with these hyperparameters, trained on each
3.5.1. Preprocessing training set, and subsequently tested on the corresponding test set
To prepare the data for the experimentation process, four prepro- within the cross-validation process. During training, neural networks
cessing steps are applied to each time series: missing data imputation, utilize the validation set to stop the training process when no further
splitting, standardization, and windowing.
improvement is observed for the validation efficacy in at least ten
Sensor errors are common and pose challenges in computer science
epochs.
and telecommunications. In this case, the sensors in the monitoring
As the optimization algorithm requires establishing a range for the
stations may not provide data for specific short periods, which can com-
hyperparameters, Table 2 summarizes it for each model. Furthermore,
plicate the experimentation process for some models. A data imputation
the window size (𝑊 ) used during preprocessing (see Section 3.5.1) is
technique is applied to address this issue. In this case, we employed
also optimized to determine the best amount of previous information
the MissForest [42] technique, which uses a random forest to fill the
data, considering all the non-missing features from the current moment together with the learning rate and the batch size in the cases of neural
during five iterations. A rigorous data imputation comparison is beyond networks. In all cases, the hyperparameters have been selected to cover
the scope of this work. For that reason, we do not consider additional a range which should generate simple models. Finally, a list of ranges
methods. is provided in Table 3 with all the hyperparameters relevant to the
Following data imputation, a cross-validation approach is employed experimentation. Note that the hyperparameters whose range has a
to assess the generalization performance of each model. In each cross- single value are not optimized.
validation iteration, an entire year is used as the test set, the previous Finally, as the Bayesian Search itself incorporates hyperparameters,
year is used as the validation set, and the remaining years are used as the default values have been chosen for use in general optimization
the training set. This means that, in each fold, the training size would problems. Specifically, the Upper Confidence Bound acquisition func-
consist of 16 years (≈140160 instances or ≈90%) while the validation tion has been used with a 𝜆 value of 2.576, a choice that has been
and test size would be one year each one (≈8760 instances or ≈5% each consistently shown to strike a good balance between exploration and
one). exploitation.
4
M.J. Jiménez-Navarro et al. Applied Soft Computing 157 (2024) 111504
Fig. 2. Results obtained by the Bayesian test applied to TLSTM (left) and L1 (right).
Table 3 Table 4
Other relevant hyperparameters. The average RMSE results obtained for the best models. We conducted a statistical
Hyperparameter Range comparison between the L1 and TLSTM models and marked the best results with an
asterisk to indicate that their differences are significant.
Window size [24, 72]
DT L1 FF LSTM TFF TLSTM
Horizon 24
Shift 24 Aljarafe 15.9 15.2 15.8 15.1 15.3 14.9*
Learning rate [0.0001, 0.01] Asomadilla 15.9 15.2 15.8 15.2 15.2 14.9*
Batch size [16, 128] Bermejales 17.9 17.0 17.8 17.2 17.2 16.9
Epochs 100 Ronda del Valle 18.1 18.0 18.1 17.7 17.4 17.0*
Early stopping patience 10 Torneo 14.7 13.9 14.5 14.1 13.9 13.9
𝜆 2.576
Average 16.5 15.9 16.4 15.8 15.8 15.5
3.5.3. Evaluation On the other hand, the Bermejales and Ronda del Valle datasets appear
The primary metric chosen was the Root Mean Squared Error to be the most difficult to predict, with an RMSE difference of 2 to 3
(RMSE) to evaluate the effectiveness of the model as this metric is compared to Aljarafe or Asomadilla. Interestingly, the Torneo dataset
widely used in air pollution problems. achieved a lower RMSE, suggesting that it might be the easiest area to
The final step in the workflow is to determine the optimal set predict.
of hyperparameters and models. This determination is based on the We can generally observe that more straightforward approaches
average efficacy in all years of cross-validation. significantly improve the results in almost all datasets. This suggests
that the problem is relatively simple and that using more complex
4. Results and discussion approaches hampers generalization. Overall, our TLSTM consistently
improves the simple LSTM by simply embedding the feature selection.
This section presents the results and discusses the implications The TLSTM achieved the best RMSE for all datasets, followed by
and hypotheses on the metrics obtained. Section 4.1 provides an the TFF, LSTM, and the L1 models with only a difference of 0.3–0.4 in
overview of the results for all datasets using grand averages. Section 4.2 terms of RMSE. On the other hand, DT and FF generally provide the
presents the RMSE distribution among the different experiments. Sec- worst results.
tion 4.3 provides the best hyperparameters obtained from evaluating There is a clear improvement in adding the TSL to the neural net-
our methodology. Section 4.4 analyzes the results on a yearly basis. works compared with their original versions (FF and LSTM), especially
Section 4.5 shows the error obtained daily and monthly. Section 4.6 in the FF neural network. We hypothesize that TSL adds flexibility to
presents the best and worst predictions achieved by the best model. the neural network, enabling it to adjust to simpler models. Thanks to
Section 4.7 illustrates which features were selected by the best model. our proposal, the neural network can adapt its complexity to match the
Finally, Section 4.8 contains the distribution of the number of features behavior of the problem. For example, in simpler problems that require
obtained in each experiment. fewer features, TSL may eliminate a few, whereas in more complex
scenarios, TSL may retain a greater number of features. For example,
4.1. General analysis this flexibility cannot be achieved with the L1 model, as it is only
suitable for linear problems.
Table 4 displays the average cross-validated metrics obtained after Finally, we compared the results between our best TSL combination
applying our methodology. Although datasets are sampled from sensors and the model with embedded feature selection.
with comparable conditions, the errors of the different models vary Specifically, the L1 and TLSTM models have been compared using a
considerably in some cases. It appears that the behavior of Aljarafe and Bayesian test [46] based on the findings in Section 4.4. The best result
Asomadilla is quite similar in terms of RMSE, which does not seem to with a significance level above 75% is indicated with an asterisk. Refer
be related to geographic proximity, as other areas studied are closer. to Fig. 2 for detailed comparison results.
5
M.J. Jiménez-Navarro et al. Applied Soft Computing 157 (2024) 111504
4.2. Error distribution analysis The FF and TFF models have consistently opted for the same number
of layers, typically one or two, without utilizing the maximum available
Fig. 3 illustrates the RMSE distribution of every experiment per- (three) layers in any dataset. As for the number of units, the FF
formed grouped by model and dataset. This plot may allow us to model typically ranged from 12 to 59, whereas the TFF model required
observe the difficulty in finding optimal hyperparameters. considerably fewer units, ranging from 7 to 18. The batch size and
In general, the neural networks are more robust to the selected learning rate do not show a consistent pattern among the datasets.
hyperparameters than the DT and L1 models employed within the The FF model seems to have increased its complexity in order to be
optimization space, even when the search space in the neural net- able to process all features, decreasing its efficacy as a consequence.
work is considerably greater. This robustness gives us greater confi- The dropout rate remained generally low, less than 0.3, except for the
dence in the model configuration even if it was selected without any Torneo and Bermejales datasets in FF.
hyperparameter optimization process. The LSTM and TLSTM models typically have a similar number of layers
The FF and TFF models exhibit a more significant difference in all across datasets, with differences only seen in Aljarafe, where TLSTM
datasets. However, the improvement is less significant between LSTM requires fewer layers. However, the situation changes when considering
and TLSTM, which could be explained as the LSTM includes some the number of units compared to the FF and TFF models. TLSTM
selection in their forget gates. However, including TSL appears to be generally requires more units, ranging from 36 to 48, whereas LSTM
beneficial for LSTM on average. ranges from 11 to 45. Additionally, TLSTM demands significantly lower
Finally, there seems not to be a relationship between the results batch sizes and learning rates compared to LSTM. Dropout rates in
obtained by our approach and the characteristics of the datasets shown LSTM are consistent across datasets, except for Bermejales.
in Table 1. In general, time series forecasting problems typically do not require
neural networks with many layers. Thus, it is reasonable that hyper-
4.3. Hyperparameters analysis parameter optimization identified a low number of layers and units as
optimal.
In general, for TSL applied to all neural networks, the regularization
This section presents and discusses the best hyperparameters found
term of the layer converged to a close range for all datasets, between
by our methodology. The goal is to identify existing patterns among the
0.002 and 0.01. Consequently, we can assume that a value within
different datasets.
this range could be a good initial value for datasets with similar
Table 5 summarizes the best hyperparameters for each dataset and
circumstances.
model. Note that the hyperparameters are in JSON format using the
Typically, a large window size in all datasets is necessary, often
‘‘key:value’’ format. In general, there seems to be no recognizable surpassing twice the forecasting horizon. Not all time steps may be
pattern that relates the optimal hyperparameters and the results in all used in the DT, L1, TFF, and TLSTM models, as they include a feature
datasets. selection process. This will be discussed in Section 4.7.
Regarding the DT model, it is clear that the best depth for the
datasets is around five or six, which is considerably less than the 4.4. Yearly analysis
maximum available. This supports the idea that the problem requires
simpler models, as increasing the depth would increase the probability This section presents the results for each year obtained after apply-
of overfitting. ing cross-validation to the best model. Note that the DT and FF models
The L1 model shows more diversity, with the predominant low val- have been removed for visual clarity, as they seem to be worse in all
ues ranging from 0.0001 to 0.02. For Aljarafe, Asomadilla, Bermejales, datasets than the L1, TFF, and TLSTM models.
and Torneo, a value of 0.0001 was the best, indicating a considerably Table 6 displays the average RMSE for each model in all datasets
low penalization for feature selection. However, the regularization is and years evaluated.
significantly greater for Ronda del Valle, suggesting that the model The results, as discussed in Section 4.1, indicate that the three most
should tends to select less features. promising models which embeds some feature selection are the L1, TFF,
6
M.J. Jiménez-Navarro et al. Applied Soft Computing 157 (2024) 111504
Table 5
Best hyperparameters found by the experimentation process.
Dataset Model Window Hyperparameters
DT 40 # Max depth: 7
FF 34 # Layers: 1, # Units: 19, Batch size: 48, Learning rate: 0.0068, Dropout: 0.0650
L1 38 Regularization: 0.0001
Aljarafe
LSTM 67 # Layers: 2, # Units: 12, Batch size: 90, Learning rate: 0.0059, Dropout: 0.2574
TFF 51 # Layers: 1, # Units: 16, Batch size: 95, Learning rate: 0.0008, Dropout: 0.2798, Regularization: 0.0096
TLSTM 65 # Layers: 1, # Units: 48, Batch size: 18, Learning rate: 0.0080, Dropout: 0.4003, Regularization: 0.0039
DT 45 # Max depth: 5
FF 67 # Layers: 2, # Units: 12, Batch size: 90, Learning rate: 0.0059, Dropout: 0.2574
L1 38 Regularization: 0.0001
Asomadilla
LSTM 67 # Layers: 2, # Units: 12, Batch size: 90, Learning rate: 0.0059, Dropout: 0.2574
TFF 64 # Layers: 1, # Units: 18, Batch size: 44, Learning rate: 0.0059, Dropout: 0.3719, Regularization: 0.0097
TLSTM 71 # Layers: 2, # Units: 48, Batch size: 51, Learning rate: 0.0002, Dropout: 0.3432, Regularization: 0.0075
DT 25 # Max depth: 4
FF 68 # Layers: 1, # Units: 45, Batch size: 93, Learning rate: 0.0014, Dropout: 0.4986
L1 38 Regularization: 0.0001
Bermejales
LSTM 68 # Layers: 1, # Units: 45, Batch size: 93, Learning rate: 0.0014, Dropout: 0.4987
TFF 47 # Layers: 2, # Units: 7, Batch size: 30, Learning rate: 0.0022, Dropout: 0.0096, Regularization: 0.0027
TLSTM 65 # Layers: 1, # Units: 36, Batch size: 29, Learning rate: 0.0004, Dropout: 0.0099, Regularization: 0.0024
DT 28 # Max depth: 6
FF 32 # Layers: 1, # Units: 56, Batch size: 114, Learning rate: 0.0005, Dropout: 0.4473
L1 56 Regularization: 0.0273
Rondadelvalle
LSTM 26 # Layers: 2, # Units: 11, Batch size: 112, Learning rate: 0.0015, Dropout: 0.3736
TFF 51 # Layers: 1, # Units: 16, Batch size: 95, Learning rate: 0.0008, Dropout: 0.2799, Regularization: 0.0097
TLSTM 71 # Layers: 2, # Units: 48, Batch size: 51, Learning rate: 0.0003, Dropout: 0.3433, Regularization: 0.0075
DT 45 # Max depth: 5
FF 24 # Layers: 2, # Units: 59, Batch size: 115, Learning rate: 0.0063, Dropout: 0.1786
L1 38 Regularization: 0.0001
Torneo
LSTM 26 # Layers: 2, # Units: 11, Batch size: 112, Learning rate: 0.0015, Dropout: 0.3736
TFF 51 # Layers: 1, # Units: 16, Batch size: 95, Learning rate: 0.0008, Dropout: 0.2799, Regularization: 0.0097
TLSTM 71 # Layers: 2, # Units: 48, Batch size: 51, Learning rate: 0.0003, Dropout: 0.3433, Regularization: 0.0075
Table 6
Best results obtained for each model and year in terms of RMSE (μg∕m3 ).
Aljarafe Asomadilla Bermejales Ronda del valle Torneo
L1 TFF TLSTM L1 TFF TLSTM L1 TFF TLSTM L1 TFF TLSTM L1 TFF TLSTM
2006 16.8 17.1 16.2 17.4 17.7 17.1 18.0 18.9 18.7 20.3 20.3 19.8 13.1 13.1 12.9
2007 15.5 16.4 15.6 16.4 16.3 16.2 17.7 18.3 17.9 19.9 19.6 19.4 13.0 13.1 13.0
2008 16.7 17.7 16.9 16.6 16.9 16.3 18.6 19.2 18.5 20.0 19.3 19.1 13.1 13.2 13.2
2009 16.6 16.9 16.7 16.2 15.8 15.8 18.7 18.6 18.7 18.6 18.4 18.1 13.5 13.7 14.4
2010 16.9 16.6 16.1 16.3 15.9 15.6 18.5 18.9 18.1 20.5 19.8 19.2 14.4 14.4 14.1
2011 15.0 16.0 14.8 15.8 15.9 16.2 17.0 17.6 17.4 18.2 17.8 17.6 13.6 13.9 13.7
2012 16.5 16.7 16.0 15.8 15.7 15.6 17.4 17.8 17.6 16.4 15.5 15.8 14.3 14.4 14.5
2013 14.9 14.4 14.2 14.6 14.4 14.7 17.2 17.2 16.7 17.2 16.7 16.1 13.7 13.7 13.5
2014 15.0 15.0 14.4 15.6 15.1 14.8 16.7 16.8 16.3 17.8 16.9 16.5 13.8 14.0 13.9
2015 17.2 16.1 16.0 14.9 14.9 14.8 16.4 16.6 16.7 18.1 17.1 16.5 14.2 14.0 14.4
2016 15.1 15.1 14.4 14.7 14.5 14.4 16.3 16.5 16.1 17.8 16.8 16.5 14.2 13.6 13.7
2017 15.2 15.1 14.5 14.8 15.1 15.2 17.2 17.8 17.2 19.2 18.1 17.7 14.6 14.2 14.1
2018 14.1 14.3 13.9 15.2 14.7 14.2 17.5 18.1 17.7 18.5 17.8 16.7 14.5 14.0 14.0
2019 13.7 13.8 14.0 15.1 14.6 14.6 16.9 17.3 16.8 17.0 16.8 16.4 14.8 14.5 14.3
2020 13.1 13.0 12.8 13.6 13.6 13.1 16.1 15.8 15.2 15.1 14.8 14.6 13.7 13.5 13.9
2021 12.9 12.9 13.1 14.1 13.7 14.1 16.2 15.7 15.9 16.7 16.1 15.4 14.2 14.4 14.1
2022 13.6 13.7 13.3 13 15.7 12.7 15.5 14.9 14.7 16.3 15.4 15.6 14.3 14.4 14.3
2023 15.2 15.2 14.6 13.6 13.3 13.3 14.0 13.4 14.3 16.3 15.6 15.4 14.0 14.3 14.3
and TLSTM models. Notably, including the TSL significantly improves 4.5. Error map analysis
the performance compared to FF in most cases, affirming the benefits
of feature selection and model simplification for this problem. In this section, we conduct a detailed analysis of the predictions
Overall, TLSTM achieves the best results in most datasets, followed obtained by our proposal for the year 2023, which is paramount in
by TFF. The L1 model obtains remarkable results, but the reliability validating the model’s effectiveness.
is slightly worse than TFF and TLSTM excepting the Bermejales and Fig. 4 illustrated the aggregated errors for each hourly prediction (y-
Torneo datasets. axis) for all months (x-axis). In the Asomadilla, Bermejales, and Torneo
It is crucial to emphasize that our primary interest lies in assessing datasets, the error is significantly lower in December compared to other
the model’s efficacy for the last year, as it represents the most ‘‘realis- months, with a similar trend observed for November and January. In
tic’’ scenario where only past data is used for training and future data Aljarafe and Ronda del Valle, January and December exhibit relatively
are used for evaluation. In this context, the TLSTM model outperforms low errors, but the distribution appears more scattered. Additionally, it
the L1 and TFF models in three datasets except for the Bermejales and is worth noting that errors at the beginning of each month are notably
Torneo datasets where TFF and L1 obtain the best results respectively. higher than on subsequent days.
7
M.J. Jiménez-Navarro et al. Applied Soft Computing 157 (2024) 111504
Fig. 4. RMSE aggregated for each month and hour of the day in Aljarafe, Asomadilla, Bermejales, Ronda del Valle, and Torneo.
4.6. Best/worst predictions analysis making the complete removal of any feature impractical. Notably, the
most frequently selected feature across almost all datasets was Ozone as
Fig. 5 displays each dataset’s best and worst 24-hour Ozone predic- expected. On the other hand, average temperature (TMP), wind speed
tions. Generally, as discussed earlier, the best predictions are observed (WP), and NO2 appeared to be less crucial than pollutants, which held
at the end and beginning of the year for most datasets, except for Aso- greater importance.
madilla and Ronda del Valle, where the best predictions occur in June.
Conversely, the worst predictions seems to occur around the middle
4.8. Number of features distribution analysis
of the year in Aljarafe and Torneo while the other worst predictions
appear at the beginning. It is essential to highlight the potential dangers
of make decisions considering the predictions when the predictions Fig. 7 represents the number of features selected by the models
have a considerable error, leading to underestimating ozone values and along the different experiments performed during the hyperparameter
poor recommendations. optimization.
In conclusion, while the model produces good overall results, it is The FF and LSTM models do not include a feature selection process,
crucial to exercise caution when interpreting predictions during periods which always selects all available features. The differences between the
of increased error. In epochs and hours with relatively low error rates, experiments are caused by the window size optimization, which also
the predictions appear to be accurate enough to be trusted. influences the number of input features.
The DT model has difficulty selecting a small number of features,
4.7. Selected features analysis probably when the optimization process is stuck in choosing a large
depth for the tree.
This section analyzes the features selected by our proposal to gain TLSTM, which performed the best in terms of RMSE, contains con-
insight into the behavior learned by the model and the problem. siderably more features than the L1 or TFF models. Generally, feature
Fig. 6 illustrates the selected features (y-axis) for each time step in selection for the TSL in LSTM is more challenging than for the TFF
the input window (x-axis) across all datasets. Two distinct behaviors model. Nevertheless, the TFF obtained remarkable results, which could
are observed among Aljarafe, Asomadilla, Bermejales, and Ronda del also be a good option.
Valle, in contrast to Torneo. Finally, the L1 and TFF models are the models that select fewer
The model opted for a larger window size in almost all datasets with features in all datasets, with L1 being remarkably better. However,
a sparse feature selection. The selected features were close to 50% in selecting too many features may cause the poor performance shown
all datasets, which suggests that most features in most time steps were in Section 4.1, oversimplifying the model in some cases to near zero
not beneficial for forecasting. Specifically, Asomadilla was the dataset features. Even though TSL does not selects fewer features, there seems
that required fewer features, with only 49.7% of the features selected. to be a more robust misconfiguration of the model and also obtained
Surprisingly, all features showed relevance in at least one-time step, better efficacy than L1.
8
M.J. Jiménez-Navarro et al. Applied Soft Computing 157 (2024) 111504
Fig. 5. Best and worst predictions for all datasets. Prediction is represented in blue, while ground truth is represented in orange.
5. Conclusions and future works promising results in the Ozone forecasting problem, even though the
solution might require a simpler model.
In this work, we present a practical application of TSL by enhanc- Looking ahead, we have several ideas planned further to improve
ing the previous proposal in more than 1 μg∕m3 in most datasets, the results and interpretability of the model. Firstly, instead of sharing
creating a more reliable model with more accurate predictions of a single TSL for every output in a multi-step problem, we assign an
the behavior of ozone pollution. Therefore, a promising methodology independent TSL for each specific output. This tailored approach can
has been provided for an efficient and environmentally aware traffic lead to more accurate predictions, giving more precise information
organization or users with respiratory problems. Our approach includes about the relevant features not just for the model but for every specific
an embedded feature selection method that simplifies and reduces the class that is not widely studied in the literature. The challenge here is to
problem dimensionality, thereby streamlining the model in cases where
modify the architecture to include this mechanism efficiently without
complexity is unnecessary to achieve good results. As a result, we
requiring extra layers/weights.
reduce both computational costs and improve the interpretability of the
Additionally, we consider the locality principle, where the moments
model. Our approach will allow us to shed some light on deep learning
near a specific moment should have a greater probability of inclusion
models with a simple yet effective approach that helps us determine
in the model. This localized focus may further refine the accuracy of
the relevant factors that influence the model prediction. TSL could be
applied to most time series forecasting problems, not only limited to the predictions, reducing the error and producing a smoother instant
environmental data. selection. The main challenge is to determine the optimal number of
Generally, in time series forecasting problems, less than 50% or 15% near instances of every selected instant and its grade of influence over
of inputs are relevant to obtaining good results. This sparse input infor- the near instances.
mation deviates from the typical approach but proves to be effective in Lastly, we plan to experiment with residual connections, as recent
our method. Notably, our approach not only yields remarkable results studies have shown that they can significantly enhance results. As
but also exhibits greater robustness to hyperparameters compared to observed in previous works, incorporating these connections into our
other models in the comparison. model may yield even better efficacy. The challenge in this improve-
Furthermore, TSL enables neural networks to handle a broader ment is to modify the different regularization functions to consider the
range of problems, regardless of their complexity. Notably, we obtained layers at every level.
9
M.J. Jiménez-Navarro et al. Applied Soft Computing 157 (2024) 111504
In general, our work demonstrates the effectiveness of the TSL ap- Writing – review & editing. F. Martínez-Álvarez: Data curation, In-
proach in time series forecasting, and we anticipate that these planned vestigation, Methodology, Supervision, Writing – review & editing. G.
enhancements will lead to even more promising outcomes in the future. Asencio-Cortés: Investigation, Software, Supervision, Writing – origi-
nal draft, Writing – review & editing.
CRediT authorship contribution statement
Declaration of competing interest
M.J. Jiménez-Navarro: Conceptualization, Data curation, Investi-
gation, Methodology, Visualization, Writing – original draft, Writing The authors declare that they have no known competing finan-
– review & editing. M. Martínez-Ballesteros: Investigation, Method- cial interests or personal relationships that could have appeared to
ology, Software, Supervision, Visualization, Writing – original draft, influence the work reported in this paper.
10
M.J. Jiménez-Navarro et al. Applied Soft Computing 157 (2024) 111504
Data availability [17] M. Dhouibi, A.K.B. Salem, A. Saidi, S.B. Saoud, Accelerating Deep Neural
Networks implementation: A survey, IET Comput. Digit. Tech. 15 (2) (2021)
79–96.
Data will be made available on request.
[18] R. Schwartz, J. Dodge, N.A. Smith, O. Etzioni, Green AI, Commun. ACM 63 (12)
(2020) 54–63.
Declaration of Generative AI and AI-assisted technologies in the [19] P. Dhal, C. Azad, A comprehensive survey on feature selection in the various
writing process fields of machine learning, Appl. Intell. 52 (4) (2022) 4543–4581.
[20] H. El, A. Mohamed, H. Fawzy, Time series forecasting using tree based methods,
J. Stat. Appl. Probab. 10 (2021) 229.
While preparing this work, the authors used ChatGPT to improve
[21] G.N. Rao, B.Y. Rao, T. Sravani, N. Ramakrishiah, M. Balaanand, LASSO-based
the readability and grammar checking of the manuscript. After using feature selection and naïve Bayes classifier for crime prediction and its type,
this tool, the authors reviewed and edited the content as needed and Serv. Orient. Comput. Appl. 13 (2019) 187–197.
took full responsibility for the publication’s content. [22] M.J. Jiménez-Navarro, M. Martínez-Ballesteros, F. Martínez-Álvarez, G. Asencio-
Cortés, PHILNet: A novel efficient approach for time series forecasting using deep
learning, Inform. Sci. 632 (2023) 815–832.
Acknowledgments [23] M.J. Jiménez-Navarro, M. Martínez-Ballesteros, I.S. Sousa-Brito, F. Martínez-
Álvarez, G. Asencio-Cortés, Feature-aware drop layer (FADL): A nonparametric
The authors would like to thank the Dirección General de Sosteni- neural network layer for feature selection, in: Proceedings of the Interna-
tional Conference on Soft Computing Models in Industrial and Environmental
bilidad Ambiental 𝑦 Cambio Climático, Consejería de Sostenibilidad,
Applications, 2022, pp. 557–566.
Medio Ambiente 𝑦 Economía Azul, for the data acquisition. The Junta [24] M.J. Jiménez-Navarro, M. Martínez-Ballesteros, F. Martínez-Álvarez, G. Asencio-
de Andalucía, Spain for projects PY20-00870 and UPO-138516. This Cortés, Embedded temporal feature selection for time series forecasting using
work has been supported by TED2021-131311B-C21 and TED2021- deep learning, in: Proceedings of International Work-Conference on Artificial
Neural Networks, 2023, pp. 15–26.
131311B-C22 funded by MICIU/AEI/10.13039/501100011033 and by
[25] A. Moslemi, A tutorial-based survey on feature selection: Recent advancements
the European Union NextGenerationEU/PRTR. This research has also on feature selection, Eng. Appl. Artif. Intell. 126 (2023) 107136.
been supported by the grant PID2020-117954RB-C22 and PID2020- [26] A. Khemphila, V. Boonjing, Heart disease classification using neural network
117954RB-C21 funded by MICIU/AEI/10.13039/501100011033. Fund- and feature selection, in: Proceedings of the International Conference on Systems
Engineering, 2011, pp. 406–409.
ing for open access publishing: Universidad de Sevilla’s library.
[27] B. Liu, Q. Cui, T. Jiang, S. Ma, A combinational feature selection and en-
semble neural network method for classification of gene expression data, BMC
References Bioinformatics 5 (2004) 136.
[28] W. Yan, Stock index futures price prediction using feature selection and deep
[1] WHO, Healthy Environments for Healthier Populations: why Do They Matter, learning, North Am. J. Econ. Finance 64 (2023) 101867.
and What Can We Do?, World Health Organization, Geneva, 2019. [29] T. Niu, J. Li, W. Wei, H. Yue, A hybrid deep learning framework integrating
[2] C.S. Malley, D.K. Henze, J.C. Kuylenstierna, H.W. Vallack, Y. Davila, S.C. feature selection and transfer learning for multi-step global horizontal irradiation
Anenberg, et al., Updated global estimates of respiratory mortality in adults ≥ 30 forecasting, Appl. Energy 326 (2022) 119964.
years of age attributable to long-term ozone exposure, Environ. Health Perspect. [30] A. Verikas, M. Bacauskiene, Feature selection with neural networks, Pattern
125 (2017) 087021. Recognit. Lett. 23 (11) (2002) 1323–1335.
[3] T. Wang, L. Xue, P. Brimblecombe, Y.F. Lam, L. Li, L. Zhang, Ozone pollu- [31] M.M. Kabir, M.M. Islam, K. Murase, A new wrapper feature selection approach
tion in China: a review of concentrations, meteorological influences, chemical using neural network, Neurocomputing 73 (16) (2010) 3273–3283.
precursors, and effects, Sci. Total Environ. 575 (2017) 1582–1596. [32] D. Tong, R. Mintram, Genetic Algorithm-Neural Network (GANN): A study
[4] A.P.K. Tai, M.V. Martin, C.L. Heald, Threat to future global food security from of neural network activation functions and depth of genetic algorithm search
climate change and ozone air pollution, Nature Clim. Change 4 (2014) 817–821. applied to feature selection, Int. J. Mach. Learn. Cybern. 1 (2010) 75–87.
[5] B. Kura, S. Verma, E. Ajdari, A. Iyer, Growing public health concerns from poor [33] A. Alshammari, Generation forecasting employing Deep Recurrent Neural Net-
urban air quality: strategies for sustainable urban living, Comput. Water Energy work with metaheruistic feature selection methodology for renewable energy
Environ. Eng. 2 (2013) 1–9. power plants, Sustain. Energy Technol. Assess. 55 (2023) 102968.
[6] K.A. Sanchez, M. Fosterc, M.J. Nieuwenhuijsend, A.D. May, T. Ramania, J. [34] B. Cancela, V. Bolón-Canedo, A. Alonso-Betanzos, E2E-FS: an end-to-end feature
Zietsmana, H. Khreisa, Urban policy interventions to reduce traffic emissions and selection method for neural networks, Clin. Orthop. Related Res. (2020).
traffic-related air T pollution: Protocol for a systematic evidence map, Environ. [35] N.L. da Costa, M.D. de Lima, R. Barbosa, Analysis and improvements on feature
Int. 142 (2020) 105826. selection methods based on artificial neural network weights, Appl. Soft Comput.
[7] X. Querol, A. Alastuey, G. Gangoiti, N. Pérez, H.K. Lee, H.R. Eun, Y. Park, E. 127 (2022) 109395.
Mantilla, M. Escudero, G. Titos, Phenomenology of summer ozone episodes over [36] D. Yuan, J. Jiang, Z. Gong, C. Nie, Y. Sun, Moldy peanuts identification based on
the Madrid Metropolitan Area, central Spain, Atmos. Chem. Phys. 18 (2018) hyperspectral images and point-centered convolutional neural network combined
6511–6533. with embedded feature selection, Comput. Electron. Agric. 197 (2022) 106963.
[8] P. Sicard, E. Paoletti, E. Agathokleous, V. Araminiené, C. Proietti, F. Coulibaly, [37] H. Zhang, J. Wang, Z. Sun, J.M. Zurada, N.R. Pal, Feature selection for neural
A.D. Marco, Ozone weekend effect in cities: Deep insights for urban air pollution networks using group lasso regularization, IEEE Trans. Knowl. Data Eng. 32 (4)
control, Environ. Res. 191 (2020) 110193. (2020) 659–673.
[9] D. Marvin, L. Nespoli, D. Strepparava, V. Medici, A data-driven approach to [38] S. Hochreiter, The vanishing gradient problem during learning recurrent neural
forecasting ground-level ozone concentration, Int. J. Forecast. 38 (3) (2022) nets and problem solutions, Int. J. Uncertain. Fuzziness Knowl.-Based Syst. 6
970–987. (1998) 107–116.
[10] D.A. Wood, Ozone air concentration trend attributes assist hours-ahead forecasts [39] P. Lara-Benítez, M. Carranza-García, J.C. Riquelme, An experimental review on
from univariate recorded data avoiding exogenous data inputs, Urban Climate deep learning for time series forecasting, Int. J. Neural Syst. 31 (2021) 2130001.
47 (2023) 101382. [40] J.F. Torres, D. Hadjout, A. Sebaa, F. Martínez-Álvarez, A. Troncoso, Deep learning
[11] K. Ko, S. Cho, R.R. Rao, Machine-learning-based near-surface ozone forecasting for time series forecasting: A survey, Big Data 9 (1) (2021) 3–21.
model with planetary boundary layer information, Sensors 22 (20) (2022). [41] European Parliament and of the Council on ambient air quality and cleaner air
[12] A. Gómez-Losada, G. Asencio-Cortés, F. Martínez-Álvarez, J.C. Riquelme, A for Europe, Directive 2008/50/EC, 2008.
novel approach to forecast urban surface-level ozone considering heterogeneous [42] D.J. Stekhoven, P. Bühlmann, MissForest—non-parametric missing value
locations and limited information, Environ. Model. Softw. 110 (2018) 52–61. imputation for mixed-type data, Bioinformatics 28 (1) (2011) 112–118.
[13] M. Martínez-Ballesteros, F. Martínez-Álvarez, A. Troncoso, J.C. Riquelme, Mining [43] J. Snoek, H. Larochelle, R.P. Adams, Practical Bayesian optimization of machine
quantitative association rules based on evolutionary computation and its ap- learning algorithms, in: Proceedings of the Advances in Neural Information
plication to atmospheric pollution, Integr. Comput.-Aided Eng. 17 (3) (2010) Processing Systems, 2012, pp. 1723–1731.
227–242. [44] D. Hadjout, A. Sebaa, J.F. Torres, F. Martínez-Álvarez, Electricity consumption
[14] M. Martínez-Ballesteros, S. Salcedo-Sanz, J.C. Riquelme, C. Casanova-Mateo, J.L. forecasting with outliers handling based on clustering and deep learning with
Camacho, Evolutionary association rules for total ozone content modeling from application to the Algerian market, Expert Syst. Appl. 227 (2023) 120123.
satellite observations, Chemometr. Intell. Lab. Syst. 109 (2) (2011) 217–227. [45] E.T. Habtemariam, K. Kekeba, M. Martínez-Ballesteros, F. Martínez-Álvarez, A
[15] M. Martínez-Ballesteros, F. Martínez-Álvarez, A. Troncoso, J.C. Riquelme, An evo- Bayesian optimization-based LSTM model for wind power forecasting in the
lutionary algorithm to discover quantitative association rules in multidimensional Adama District, Ethiopia, Energies 16 (5) (2023) 2317.
time series, Soft Comput. 15 (2011) 2065–2084. [46] A. Benavoli, G. Corani, J. Demsar, M. Zaffalon, Time for a change: a tutorial for
[16] M.J. Jiménez-Navarro, M. Martínez-Ballesteros, F. Martínez-Álvarez, G. Asencio- comparing multiple classifiers through Bayesian analysis, J. Mach. Learn. Res.
Cortés, A new deep learning architecture with inductive bias balance for 18 (2017) 2653–2688.
transformer oil temperature forecasting, J. Big Data 10 (1) (2023) 80.
11