0% found this document useful (0 votes)
7 views16 pages

Data-Driven Approach To Spatial

This study presents a data-driven approach to spatial drought mapping using machine learning techniques to create drought vulnerability maps (DVMs) for Pakistan. It introduces a Combined Drought Index (CDI) based on a weighted Entropy-based TOPSIS method and evaluates various ensemble machine learning models, with Support Vector Machine (SVM) showing the highest accuracy. The research emphasizes the importance of effective drought management strategies to mitigate socioeconomic impacts and identifies highly affected regions.

Uploaded by

mariam90.ais
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views16 pages

Data-Driven Approach To Spatial

This study presents a data-driven approach to spatial drought mapping using machine learning techniques to create drought vulnerability maps (DVMs) for Pakistan. It introduces a Combined Drought Index (CDI) based on a weighted Entropy-based TOPSIS method and evaluates various ensemble machine learning models, with Support Vector Machine (SVM) showing the highest accuracy. The research emphasizes the importance of effective drought management strategies to mitigate socioeconomic impacts and identifies highly affected regions.

Uploaded by

mariam90.ais
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Earth Systems and Environment

[Link]

ORIGINAL ARTICLE

A Data-Driven Approach to Spatial Drought Mapping Using Machine


Learning
Farheen Fatima1 · Ijaz Hussain1,3 · Hanen Louati2 · Jianyi Lin3 · Ibrahim A. Nafisah4 · Mohammed M. A. Almazah5

Received: 11 December 2024 / Revised: 9 April 2025 / Accepted: 11 April 2025


© King Abdulaziz University and Springer Nature Switzerland AG 2025

Abstract
Drought is a critical barrier to socioeconomic development, necessitating effective modeling to mitigate its impacts.
Drought vulnerability modeling is crucial for managing and lessening the effects of drought. Creating a drought vulner-
ability map is the first step in developing a comprehensive drought management strategy, which is crucial for reducing
the likelihood of droughts. This study introduces a novel Combined Drought Index (CDI) developed through a weighted
Entropy-based TOPSIS method to facilitate comprehensive drought analysis. Building on this, we employed a suite of
ensemble machine learning techniques, including M5P, Dagging, Random SubSpace (RSS), Random Forest (RF), and
Support Vector Machine (SVM) models, to evaluate drought vulnerability maps (DVMs) across Pakistan. The models
were trained on 70% of the data, with the remaining 30% reserved for validation. Performance was assessed using RMSE,
MSE, MAE, and R-square metrics. Among the models, the SVM exhibited the highest accuracy in capturing drought
vulnerability, suggesting its potential for reliable spatial assessment.

Highlights
● Developing a comprehensive drought management strategy is crucial for drought monitoring.
● Propose a combined drought index based on twelve factors, whose weights were defined using an entropy-based TOPSIS
approach.
● Ensemble machine learning techniques, including M5P, Dagging, Random SubSpace (RSS), Random Forest (RF), and
Support Vector Machine (SVM) models.
● Identification of regions that are highly affected by drought.
● SVM exhibited the highest accuracy in capturing drought vulnerability.
● Graphical Abstract Description: Creating a drought vulnerability map and developing a comprehensive drought man-
agement strategy is crucial for reducing the likelihood of droughts. We propose a combined drought index based on
twelve factors, whose weights were defined using an entropy-based TOPSIS approach. Drought vulnerability maps are
created using kriging and machine learning techniques. It is concluded that the support vector machine outperforms other
methods.

3
Ijaz Hussain State Key Laboratory for Ecological Security of Regions and
ijaz@[Link] Cities, Institute of Urban Environment, Chinese Academy of
Sciences, Beijing, China
1
Department of Statistics, Quaid-i-Azam University, 4
Department of Statistics and Operations Research, College of
Islamabad, Pakistan
Sciences, King Saud University, P. O. BOX 2454,
2
Mathematics Department, Faculty of Science, Northern Riyadh 11451, Saudi Arabia
Border University, Arar, KSA, Saudi Arabia 5
Department of Mathematics, College of Sciences and Arts
(Muhyil), King Khalid University, Muhyil
61421, Saudi Arabia

13
F. Fatima et al.

Graphical Abstract

Keywords Combined drought index · Drought vulnerability mapping · Entropy · GIS · Machine learning models

1 Introduction Droughts have a significant effect on socio-economic,


environmental, agricultural, and economic activities, as
Drought is a natural hazard. Natural hazards are occur- well as the availability of water. Meteorologists catego-
ring more frequently and with greater intensity and sever- rize drought into various classifications, including meteo-
ity. People must develop stronger resilience to manage rological, socio-economic, agricultural, and hydrological
risks and adapt to these changing environmental conditions drought. A meteorological drought is diagnosed by measur-
effectively. Variations in hydrometeorological factors and ing a longer-term deficit in rainfall; an agricultural drought
economic variables, coupled with the unpredictable nature is diagnosed by measuring a shortfall in soil moisture,
of water demands in various regions globally, have made groundwater, or rainfall. A hydrological drought occurs
it challenging to define drought precisely. Yevjevich et al. when there is insufficient surface water available. The final
(1967) noted that the diverse perspectives on drought defini- stage of drought, known as the socio-economic drought, is
tions are significant barriers to drought research.

13
A Data-Driven Approach to Spatial Drought Mapping Using Machine Learning

characterized by a persistent and long-term shortage in agri- development (Cheng et al. 2016; Wu et al. 2018; Shen et al.
cultural yield and output volume. 2020; and Tang et al. 2021) as well as the regular advance-
Climate change has intensified the severity and fre- ment of GIS information (Moayedi et al. 2019; Nguyen et
quency of hydrometeorological disasters in Pakistan (Shaw al. 2019) and satellite imaging (Zuo et al. 2019; Zhang et al.
et al., 2015). Historically, the country has faced numerous 2021). Numerous analytical techniques have been proposed
severe droughts, which have significantly impacted agricul- to address scientific issues (Trajković et al. 2020; Wable et
tural productivity and the overall GDP (Ashraf and Routray al. 2019; Tian et al. 2018). Numerous studies have utilized
et al., 2015). Although drought events have been less fre- various machine learning techniques to predict or forecast
quent than floods, Pakistan has experienced several signifi- drought conditions globally. Methods such as artificial neu-
cant droughts throughout its history. For instance, Punjab ral networks (Nabipour et al. 2020), random forests (Dikshit
suffered severe droughts in 1899, 1920, and 1935. Khyber et al. 2021), and support vector machines (Zahraie et al.,
Pakhtunkhwa (KPK) encountered its worst droughts in 2011) have demonstrated promising results in hydrological
1902 and 1951. Sindh and Balochistan faced major droughts and meteorological drought forecasting.
in different years, i.e., 1871, 1881, 1899, 1931, 1947, and This study utilizes various physical factors to assess
1999. The majority of Pakistan has an arid to semi-arid cli- drought in a region. To achieve this, the Entropy-Weighted
mate, receiving less than 250 mm of rainfall annually, while method was employed, providing essential weights for
a small Northern region experiences more humid conditions each component and aligning them according to their sig-
(Khan et al., 2003). nificance. Following this, machine learning models were
The Standardized Precipitation Index (SPI), the most applied to spatially assess drought conditions across Paki-
widely used technique for determining drought world- stan. Different validation methods, including MSE, RMSE,
wide, was proposed by McKee et al. (1993) and utilizes and R-squared, were employed for model evaluation.
only monthly precipitation data. Due to its simplicity in
calculation and focus on precipitation as the sole meteo-
rological parameter, this index is widely used in practice. 2 Materials and Methods
Many multi-criteria decision-making techniques have been
employed to evaluate the available alternatives, including 2.1 Study Area
AHP, PCA, and the entropy method, when multiple param-
eters are considered for assessment. Employing geospatial Pakistan is a developing country with a long history of natu-
analysis and the Analytic Hierarchy Process (AHP), vul- ral disasters. Latitude 30.3753°N and longitude 69.3451°E
nerability to drought events in northwestern Bangladesh are the coordinates of Pakistan. The four provinces that
was found to be moderate to extreme in 77% of the region make up Pakistan are Sindh, Punjab, Khyber Pakhtunkhwa
(Hoque et al. 2021). In Bangladesh, GIS-based assessments (KPK), and Balochistan. The monsoon season brings most
have revealed that the northwestern and southwestern parts of the yearly precipitation to Sindh and Punjab, while the
are particularly vulnerable to agricultural droughts. How- winter months provide virtually little. The most incredible
ever, the central region also experiences extreme drought province, Balochistan, receives a wide variety of annual
conditions (Aziz et al. 2022). The Technique for Order of precipitation, with summer and winter rainfall ranging from
Preference by Similarity to Ideal Solution (TOPSIS), as pro- 112 to 402 mm to 71–231 mm, respectively. Yearly rainfall
posed by Hwang et al. (1981), was employed in this study in Khyber Pakhtunkhwa (KPK) and the Federally Admin-
to analyze the drought. It can be used to evaluate supply istered Tribal Areas (FATA) ranges from 250 to 1,450 mm,
chain, manufacturing process, financial and organizational with greater totals in the northern areas. Due to its high
performance, customer-focused product design processes, altitude, Gilgit-Baltistan (GB) experiences frigid tempera-
multipurpose inventory planning, risk assessment, data tures, primarily influenced by disturbances from the west,
mining, supplier, and facility location selection (Gokkaya et which result in precipitation in the area. For this study, data
al., 2017). Babaei et al. (2013) and Hadisuwito et al. (2018) on various physical factors of drought were obtained from
proposed hydrological studies related to drought that utilize NASA’s website for meteorological stations in Khyber Pak-
the TOPSIS technique. Topcu et al. (2022) employed this htunkhwa and Punjab (see Fig. 1) from 1981 to 2022.
method to conduct a drought analysis at Kars station, Tur-
key, using various meteorological parameters. 2.2 Methodology
The integration and analysis of data points from many
sources are made more accessible by a spatial examina- Here, we discuss the methods we have applied for the spa-
tion tool (Palchaudhuri et al., 2016). Subsequently, there tial assessment of drought. In the subsequent five phases,
were notable developments in the field of innovation well-known MLAs created the drought vulnerability maps.

13
F. Fatima et al.

Fig. 1 This map visually represents the study area, including specific locations of data collection points used in this study

Step 1: Drought parameter selection: The choice of (Table 1). For drought vulnerability mapping, it is essential
drought vulnerability parameters depended on a literature to identify current drought-affected regions. Based on the
review and the current climate conditions. available data, different factor layers have been generated
Step 2: New Drought Index: Develop a new Combined within the ArcGIS environment, as shown in Fig. 3. A total
Drought Index(CDI) using the Entropy weight-based TOP- of twelve meteorological parameters were selected based
SIS method. on the study area’s geo-environmental condition and previ-
Step 3: Construction of Spatially Distributed Data Lay- ous research (Table 2). Data integration and analysis were
ers: Data from drought-affected areas and drought vulner- completed after considering all layers. The base resolution
ability factors (DVFs) were collected to predict spatial for the final drought vulnerability maps was the DEM reso-
drought. lution (30 m x 30 m). The model used as a basis classifier
Step 4: Preparation of Drought Vulnerability Maps: for some ensemble models, including Dagging and RSS, is
Ensemble and machine learning techniques (M5P, Dagging, M5P.
RSS, Random Forest, and Support Vector Machine) were The factors used in this study are obtained from differ-
employed to create drought vulnerability maps, utilizing ent studies. Region-wise meteorological data are calculated
training datasets for assistance. for different stations. The location of the meteorological
Step 5: Model validation and comparison: The MSE, stations is shown in Fig. 1. The following parameters are
MAE, R2 , and RMSE tests were used for model validation. calculated for every location:
Fig. 2.
1. Three months of extreme drought frequency.
2.3 Construction of Spatial Data Layers 2. Six months of extreme drought frequency.
3. Twelve months of extreme drought frequency.
The spatial distribution of drought characteristics derived 4. Twenty-four months of extreme drought frequency.
from the SPI series was visualized using the Kriging method 5. Three months extreme drought return period.

13
A Data-Driven Approach to Spatial Drought Mapping Using Machine Learning

Fig. 2 Schematic representation of the proposed methodology

Table 1 Classification of drought according to SPI values distribution functions are not defined at x = 0. So, the
Values Drought Classes H (x) can be calculated as:
More than 0 Non-Drought
0 to −1.0 Mild Drought H (x) = q + (1 − q) · G (x)
−1.0 to −1.5 Moderate Drought
−1.5 to −2.0 Severe Drought where q represents the probability of zero and G (x) repre-
Less than − 2 Extreme Drought sents the CDF of the distribution function.
McKee et al. (1993) classified SPI values for classifying
6. Six months extreme drought return period. droughts.
7. Twelve months extreme drought return period. The severity of drought incidence was only assessed
8. Twenty-four months extreme drought return period. using the extreme drought classification. The formula used
9. Annual Rainfall. for drought frequency is as follows:
10. Rainfall Trend.
11. Temperature Trend. Ni
DFi,100 = × 100
12. Annual Humidity. i· n

2.3.1 Standardized Precipitation Index (SPI) where the frequency of droughts for time scale i (3, 6, 12,
and 24 months) over 100 years is shown by DFi,100 . Dur-
The SPI values were calculated to assess the severity of the ing every given time scale i, the number of drought months
drought (McKee et al. 1993). Calculating the Standardized during the n-year period is represented by Ni , where i is
Precipitation Index (SPI) requires only precipitation data. the time scale (e.g., 3,6, 12, and 24 months).
It may be used to evaluate the state of the drought for vari- Return Period: The California method (Wable et al. 2019)
ous periods, including 48, 24, 12, 6, 3, and 1 months, as was utilized to determine the Return Period (RP) of extreme
calculated using this rainfall data (Mehr et al. 2020). It can droughts. This involved sorting all SPI values in ascending
be used to define and compare drought conditions in differ- order and assigning ranks to each SPI value.
ent areas. Considering that the gamma function and other

13
F. Fatima et al.

Fig. 3 Factors of physical drought: A Drought frequency for 3-mon, for 6-mon drought frequency, G RI for 12-mon drought frequency, H
B Drought frequency for 6-mon, C Drought frequency for 12-mon, D RI for 24-mon drought frequency, I Annual rainfall, J Annual humid-
Drought frequency for 24-mon, E RI for 3-mon extreme drought, F RI ity, K Rainfall trend, L Temperature trend

13
A Data-Driven Approach to Spatial Drought Mapping Using Machine Learning

Table 2 Variables and their directionality of influence on drought vul- 1. At first, create the decision matrix where rows repre-
nerability
sent alternatives and columns represent criteria.
Sl Variables Direction-
no. ality of    
x11 ··· x1j ··· x1n D1 (xi )
Influence  .. .. .. .. ..   .. 
1 Rainfall Inverse  . . . . .   . 
   
D= xi1 ··· xij ··· xin = Di (xj ) (1)
2 Rainfall trend Inverse  .. .. .. .. ..   .. 
 . . . . .   . 
3 Extreme drought frequency for 24 months(%) Direct
xm1 ··· xmj ··· xmn Dm (xn )
4 Extreme drought frequency for 12 months (%) Direct
5 Extreme drought frequency for 6 months(%) Direct 2. In order to construct a standard matrix based on the nor-
6 Extreme drought frequency for three months(%) Direct
malized vector rij , the feature matrix is normalized. Equ 2
7 The return period of 24 months of extreme Inverse
drought
is used to calculate the normalized vector r. ij . Importantly,
8 The return period of 12 months of extreme Inverse there is no significance in the computation of logarithms of
drought zero and negative values. On the other hand, some values of
9 Return period of 6 months extreme drought Inverse climatic data, similar to temperature, may have zero values
10 The return period of 3 months of extreme Inverse in the case of rainfall or negative values otherwise.
drought
11 Humidity Inverse rij = max(x
ij x −min(x )
ij
,
12 Temperature trend Direct ij )−min(xij ) (2)
i = 1,2, . . . , m; j = 1,2, . . . , n
n
RI = 3. The process of calculating the weight of every variable by
p
using the entropy method is as follows:
Annual precipitation: As precipitation increases, drought 1) Finding the proportion pij of the project i index value
will decrease (Antwi-Agyei et al. 2012). The sum of the under the index j : pij is computed as in Eq. 3:
monthly totals has been used to get the yearly average of
rij
each district’s monthly precipitation. pij = ∑m (3)
Rainfall and Temperature Trend: The Mann-Kendall test i=1 rij

(Mann et al., 1945) was used to assess trends for rainfall


and temperature. A downward trend in rainfall indicates 2) Determine the index’s entropy, or ej as specified in Eq.
worsening dry conditions, while an upward trend suggests 4:
improvement. Recognizing dry spell weakness, temperature ∑m
is a significant parameter (Liang et al. 2014). A region’s ej = −k i=1 pij ln (pij ) (4)
vulnerability and dryness will increase as temperature rises.
The formula for calculating this trend is as follows: k in Eq. (4) can be calculated as:
1
n−1
∑ n
∑ k= ln(m) (5)
S= sgn(xj − xk )
k=1 j=k+1
3) Calculate the entropy weight wj of the index, the for-
mula given in Eq. 7:
2.4 New Combined Drought Index ∑n
(1−ej )
wj = ∑n (1−e )
, j=1 wj = 1 (6)
In the study of information theory, entropy is frequently
j
j=1

used to measure the degree of information disorder (Lin et


al. 2008). The entropy weight is utilized to account for vari- 4) Based on the weight’s standardized value vij , that is cal-
ations in records across several schemes. The more valuable culated by using the formula:
the data is for a final choice, and the larger the disparity
between the records under various situations, the bigger the vij = wj rij  (7)
entropy weight is. Conversely, the distinction of the record
in various plans is less. The objective data’s information is find the ideal solution A* and the anti-ideal solution A−
reflected in entropy. The objectivity reflects how robust the . The locations of A* and A− are found in Eqs. 8 and 9:
data given by the record in the assessment is.
The following is a step-by-step guide to estimating the A* = {( maxvij |j ∈ J), (minvij | j ∈ J′)}
 (8)
proposed CDI using this method: = {v1* , v2* , . . . , vn* }

13
F. Fatima et al.

A− = {( minvij |j ∈ J), (maxvij | j ∈ J′)} the better the evaluation objective is. Table 3 shows the clas-
(9)
= {v1− , v2− , . . . , vn− } sification of drought on the basis of CDI:

In Eqs. (8) and (9), J1 is the set of indices j that are profit- 2.5 Machine Learning Models for Drought Mapping
ability indices, representing the optimum values. Similarly,
J2 is the set of indices j that are loss indices, representing 2.5.1 M5P
the worst values.
The distance between index j and the best objective Decision trees are commonly utilized in machine learning
is represented by vj* in these equations, and the distance and data analysis fields to provide a clear and visual rep-
between index j and the worst objective is represented resentation of decision-making processes for classification
by vj− The evaluation result’s performance improves with and regression. M5P is a reproduced model of Quinlan’s
M5 calculation (Quinlan et al. 1992) for building regression
larger values of profitability indices or smaller values of loss
model trees. Using multifactor linear methods, this approach
indices.
generates trees. It is more efficient and accurate to produce
5) Distance scale, defined by Euclidean distance, is the results when smaller trees are used. The M5 tree technique
measure of the separation between each objective and the manages continuous class issues rather than discrete classes
ideal solution or against the ideal solution. The distance and can deal with assignments with high dimensionality. A
between the goal and the ideal solution A* is represented by detailed analysis can be found in (Talukdar et al. 2020).
S*, while the distance between the goal and the anti-ideal The data about the dividing rules for the M5 model tree
solution A- is represented by S-. The formula is defined in is acquired based on computations of error at every node.
Eqs. 10 and 11: The class standard deviation examines the mistake esteems
√∑ that show up at a node. For splitting at the node, the attribute
2 with the most significant expected error reduction from test-
S* = j=1 (vij − vj* )  (10)
n

ing each attribute is chosen. The SDR is calculated as:


√∑
2
S− = j=1 (vij − vj− ) (11) ∑ |Ki |
n
SDR = sd (K) − sd (Ki )
|K|
here, i = 1,2,…,m, and S* indicates how near the desired
goal each evaluation is. The program is more preferable; the where K represents the set of instances that reach the node;
smaller S* value indicates the shorter distance between the Ki denotes the subset of instances that have the i-th value
aim and perfect solution. of the possible set, and sd stands for the standard deviation.
6) Computing the ideal solution’s closeness degree C*, Figure 4 shows the flowchart of the M5P model:
defined in the formula given in Eq. 12
2.5.2 Dagging
Si−
Ci* = Si− +Si*
(12)
The dagging algorithm, also known as disjoint aggregating,
is another method of Ensemble Machine Learning (EML)
Ci* is in the range of 0 to 1. Ai is the most optimal evalu- employed to develop meta-learners (Zounemat-Kermani et
ation objective when Ci∗ = 0, so Ai = A*. The value of al. 2021). Dagging and bagging are similar, yet the testing
Ci∗ is used to group all of the evaluation goals, which range method is unique. This method utilizes the disjoint sam-
in size from small to large. The bigger the worth of Ci∗ is, pling method instead of bootstrap sampling to find random-
ized extracting segments from the original dataset without
Table 3 Classification of drought according to CDI values replacement (Barzegar et al. 2021). This meta-learner serves
Drought Category Interval each piece of data as a copy of the provided base learner
Extremely Wet 0.9 < TOPSIS < 1 and creates several disjointed, layered folds from the data.
Severe Wet 0.8 < TOPSIS < 0.9 Instead of bootstrap samples, it uses multiple disjoint sam-
Medium Wet 0.7 < TOPSIS < 0.8 ples to derive the base learner. According to a computational
Weak Wet 0.6 < TOPSIS < 0.7 perspective, with N designs comprising the preparation
Normal 0.4 < TOPSIS < 0.6 dataset, dagging builds M information subsets, each with
Weak Drought 0.3 < TOPSIS < 0.4 n designs, but without repeating any individual example.
Medium Drought 0.2 < TOPSIS < 0.3
As a result, a distinct model and predictions are developed
Severe Drought 0.1 < TOPSIS < 0.2
Extremely Drought 0 < TOPSIS < 0.1

13
A Data-Driven Approach to Spatial Drought Mapping Using Machine Learning

The final result is obtained by aggregating the results of the


classifiers operating in parallel (Dong et al. 2020). The num-
ber of seeds and iterations, which are usually tuned by trial
and error, are important factors for this model (Mosavi et al.
2020). The Random Subspace Method performs best when
discriminative data is dispersed over many characteristics.
However, the random subspace technique usually performs
poorly when features are less informative and data is noisy
(Sammut et al., 2017). Figure 4 displays the RSS model’s
block diagram.

2.5.4 Random Forest

It is an ensemble classifier that utilizes numerous decision


tree techniques for both classification and regression tasks
(Svetnik et al. 2003). It has a quick learning curve and pro-
duces very accurate classifiers. Without the need for variable
selection, RF can handle thousands of input variables and
performs well on massive datasets. Additionally, it is effec-
tive in detecting interactions among variables. It is effective
in detecting interactions between variables. It is composed
of an ensemble of tree-structured classifiers, denoted as
fh (x, Θ K ) for K = 1, . . . , n, where Θ K represents
Fig. 4 Flowchart for building a decision tree with linear regression
independent, identically distributed random trees. Each tree
functions (Meshram et al. 2023)
contributes a single vote towards the final classification of
for each dataset; the final model is then constructed using the input x. Similar to the Classification and Regression
the mean (average) of their predictions (Ting et al., 1997). Tree (CART) algorithm, RF utilizes the Gini index to deter-
mine the final class in each decision unit. The final classifier
2.5.3 Random Subspace is constructed by aggregating and voting the final class of
each tree using weighted values. The selection of a random
The Random Subspace Method (RSS), first introduced in seed preserves the class distribution while randomly select-
1988, improves the performance of individual models and ing a subset of samples from the training dataset. A set of
increases the reliability of less robust classifiers (Pham et al. RF attributes from the original dataset is selected using the
2018). Using random sampling, RSS generates many feature chosen database, considering user-defined values Fig. 5.
subspaces and uses these subspaces to train base classifiers.
Fig. 5 Flowchart of the ensemble
method for the RSS model (Elbel-
tagi et al. 2023)

13
F. Fatima et al.

2.5.5 Support Vector Machine drought frequency has the most excellent weight effect.
Table 4 lists the climatic parameters’ entropy weights.
SVM can be categorized in “supervised learning methods”, By applying CDI in selected regions, Nowshera,
which are based on statistical learning theory and the princi- Charsadda, Mardan, and Cherat were shown as the wettest
ple of structural risk minimization (Roodposhti et al. 2017). regions. At the same time, Rahim Yar Khan, D G Khan,
An optimal hyperplane classifies new examples effectively. and Rajanpur were the driest regions of the selected loca-
It involves several key tuning parameters: Kernel, Regular- tions. The detailed output of CDI display is shown in Fig.
ization, Gamma, and Margin. In SVM, the margin refers to 8, which shows the classification according to the observed
the perpendicular distance between the nearest data points values of CDI. The highest value ( > 0.40) shows the wet-
and the hyperplane (Arabameri et al. 2020). Regularization test conditions, and the lowest values ( < 0.40) show the
helps fine-tune the SVM classifier to minimize the mis- driest conditions.
classification of both small and large data points. Gamma When the TOPSIS method was applied to different mete-
determines the calculation for the possible separation line. orological stations in the data, it was observed that three sta-
The Kernel is a mathematical function used to transform the tions showed arid conditions: thirteen were severely dry, ten
data through linear algebra, with various SVM algorithms were moderately dry, twelve were mildly dry, fifteen were in
employing different kernel functions. SVM utilizes several the category of near to normal, and seven locations showed
types of kernels, including linear, non-linear, radial basis above-normal conditions. The frequency of drought condi-
function, sigmoid, polynomial, and exponential, each ensur- tions is demonstrated in Fig. 9.
ing a distinct margin of separation Fig. 6. After creating a new index, the multicollinearity of the
variables was assessed using the variance inflation factor
and tolerance value limits. Based on these assessments,
3 Results and Discussion twelve factors were selected for computing drought vulner-
ability in the region.
3.1 New Combined Drought Index The ranges of VIF and TOL are 1.26 to 2.0 and 0.499 to
0.793, respectively, as shown in Table 5. It shows a high
The weights of the study’s parameters were determined by VIF value of temperature trend, which is 2.0009, and the
applying the Entropy method, which is a component of the lowest value of VIF is shown for variable 12-month extreme
TOPSIS method. In this work, twelve distinct environmen- drought frequency. Similarly, the above table shows TOL
tal factors were used as TOPSIS technique input parameters values for different variables, where the smallest value of
to analyze drought for different meteorological stations Fig. 0.5111 is for average rainfall, and the highest value, 0.7930,
7. The findings show that mean annual relative humidity is for the 12-month extreme drought frequency variable
has the least weight effect, and twenty-four-month extreme
Fig. 6 Steps involved in solving a
classification/regression problem
using random forest (Achite et al.
2023)

13
A Data-Driven Approach to Spatial Drought Mapping Using Machine Learning

Fig. 7 The Block diagram of the support vector machine illustrating the key stages of the model (Markuna et al. 2023)

Table 4 Weight calculation for different factors using entropy methods


for meteorological stations
Parameter Entropy
Weights
Extreme drought frequency for three months 0.1030
Extreme drought frequency for six months 0.0714
Extreme drought frequency for twelve months 0.1530
Extreme drought frequency for twenty-four months 0.1876
The return period of three months extreme drought 0.0707
The return period of six months of extreme drought 0.0674
The return period of twelve months extreme drought 0.1263
Return period of twenty-four months extreme drought 0.1641
Annual Rainfall 0.0067
Rainfall Trend 0.0240
Temperature Trend 0.0216
Annual Humidity 0.0035
Total entropy weights 1

Table 6. Thus, there is no problem of multicollinearity


between parameters.

3.2 Mapping Using Machine Learning Models

The Machine learning methods were used to create drought


vulnerability maps (Fig. 10). The expected drought map for
each model was split into five different classes using the
natural-break method (Hoque et al. 2021). The M5P model
covers 31.91% and 21.9014% of the state in high and very
high exposure zones. Except for these parts of the area where
moderate, low, and very low zones accounted for 18.721%, Fig. 8 Spatial classification of combined drought index (CDI) values
by location
15.241%, and 12.224% of the entire area, respectively. For
the Dagging model, very low and low droughts are exposed
to 32.9981% of the whole area. Very high and high zones
are composed of 50.0893% of the area. Moderate zone

13
F. Fatima et al.

Fig. 9 Frequency of drought events


according to drought classification

Table 5 Multicollinearity results using variance inflation factor and Table 6 Drought assessment area for different models
tolerance values Models Class Area (%) Area ([Link])
Variables VIF TOL M5P Very Low 12.224 34588.8
Extreme drought frequency for three months 1.6898 0.5917 Low 15.241 43125.87
Extreme drought frequency for six months 1.6646 0.6007 Moderate 18.721 52971.2
Extreme drought frequency for twelve months 1.2609 0.7930 High 31.91 90,290
Extreme drought frequency for twenty-four 1.3952 0.7167 Very High 21.901 61968.83
months Dagging Very Low 16.650 47110.728
Return period of three months extreme drought 1.6535 0.6047 Low 16.348 46256.354
Return period of six months extreme drought 1.4774 0.6768 Moderate 16.912 47851.631
Return period of twelve months extreme drought 1.2775 0.7827 High 27.282 77194.045
Return period of twenty-four months extreme 1.4752 0.6778 Very High 22.807 64531.952
drought
RSS Very Low 16.560 46857.086
Rainfall Trend 1.4907 0.6708
Low 16.503 46696.891
Average Rainfall 1.9564 0.5111
Moderate 17.959 50815.241
Temperature Trend 2.0009 0.4997
High 27.138 76786.883
Average Humidity 1.7922 0.5579
Very High 21.837 61788.610
RF Very Low 16.072 45475.402
occupied 16.912% area. RSS model comprised 48.976% of Low 17.249 48806.127
the model for high and very high class. Moderate drought Moderate 16.072 45475.402
captured 17.959% of the total area. The low and very low High 34.786 98426.580
classes comprised 16.503% and 16.560% of the area, respec- Very High 15.819 44761.199
tively. For the Random Forest model, very low, low, and SVM Very Low 11.882 33620.960
moderate drought zones accounted for 16.072%, 17.249%, Low 13.776 38980.823
and 16.072% of the grand total area, respectively. Very Moderate 17.558 49680.526
high and high zones captured by 143187.7798 [Link] of the High 25.076 70953.109
area. SVM model shows very low and low drought zones Very High 31.705 89709.293
in 72,601.7845 [Link] of the total land. Moderate drought
accounted for 49680.52607 [Link]. A very high and high 3.3 Comparison Among Different Models of
drought zone was captured in a 160662.4026 sq. km area. Drought

Assessing the models’ logic is a crucial step in determin-


ing their predictive ability. The MSE, RMSE, MAE, and
r-square techniques were used to validate the models.
Table 7 shows the results of different models. The
SVM model shows better results than all the others, with

13
A Data-Driven Approach to Spatial Drought Mapping Using Machine Learning

Fig. 10 Drought maps by using models: A Map by M5P, B Map by Dagging, C Map by RSS, D Map by RF, E Map by SVM

Table 7 Model performance metrics


MSE = 0.00094, RMSE = 0.0307, MAE = 0.0223, and R2
Models MSE RMSE MAE R2 = 0.9627. The RF model performs worst, with results MSE
M5P 0.0048 0.0694 0.0584 0.8270
= 0.0074, RMSE = 0.0861, MAE = 0.0659, and r-square
Dagging 0.0044 0.0666 0.0417 0.8526
= 0.8662. The four ensemble models exhibited outstand-
RSS 0.0033 0.0578 0.0463 0.9133
ing predictive abilities in generating a drought vulnerability
RF 0.0074 0.0861 0.0659 0.8662
SVM 0.0009 0.0307 0.0223 0.9627
map. So, the result shows that the SVM model is the most
appropriate fit for drought mapping.

13
F. Fatima et al.

3.4 Discussion 4 Conclusion

One of the most dangerous aspects of the climate is drought. This study has adapted and utilized different ensemble
It negatively impacts the standard of living for the major- methodologies for evaluating drought vulnerability. The
ity of people in locations where agriculture is the primary TOPSIS model was used to analyze drought data by cre-
source of income. Researchers concentrated more on ating a new drought index at different meteorological sta-
drought prediction than drought assessment in the majority tions of Khyber Pakhtunkhwa and Punjab, utilizing the
of earlier investigations. However, an evaluation of drought entropy weight method and twelve climatic parameters to
susceptibility that takes into account the many physical fac- explore the situation. The new index is classified into dif-
tors is necessary for developing scientific methods to lessen ferent classes according to classification criteria. Severe dry
the impact of drought. The current study assessed drought and near-normal drought incidents are frequently observed.
using a variety of factors, including temperature, humidity, The study creates maps of drought susceptibility using a
and precipitation. Factors were employed to account for all GIS environment and twelve drought-determining factors
potential drought conditions in order to undertake this inves- by using different machine learning methods. The M5P,
tigation. The proven method of SPI-based drought estima- Dagging, RSS, random forest, and support vector machine
tion has been used in the evaluation of drought (Malik et al. models were combined, revealing high vulnerability rates
2020; Mehr et al. 2020). of 21.901%, 22.807%, 21.837%,15.819%, and 31.705%
The criteria were selected based on the geoenvironmen- in the region. Different factor layers were created for the
tal parameters of the study region and previous research. spatial assessment of drought, and the Kriging interpolation
Well-known Machine Learning Algorithms (MLAs) were method in ArcGIS was used. Maps of different models were
used in the evaluation. Drought susceptibility was assessed divided into five breaks: very high and high, moderate, very
using five ensemble and machine learning models: M5P, low, and low. Different model validation techniques, MSE,
M5P-Dagging, M5P-RSS, RF, and SVM. The maps were RMSE, MAE, and R2 are applied to the model’s result.
created by applying each of these models—M5P, Dagging, These techniques help to identify which model performs
RSS, RF, and SVM—and by taking into account all relevant best; according to the results, the SVM model performs
aspects. In a variety of fields, including stream flow predic- best among other models. It shows which region is highly
tion (Onyari et al.,2013), flood hazard (Nhu et al. 2020b), affected by drought and which is not. The absence of effec-
landslide (Antronico et al. 2020), assessment of defores- tive drought management strategies could lead to future vul-
tation susceptibility (Saha et al. 2021) and gully erosion nerability in this area.
(Nhu et al. 2020a); Roy and (Saha et al. 2021) a number of
researchers employed M5P, Dagging, RTF and RSS MLAs. Data Availability The data that support the findings of this study are
available from the corresponding author upon reasonable request.
As in the previously listed fields, every model used in
this study produced excellent results. Among the models
that were used, SVM had the best accuracy (96.27%). M5P
Declarations
was employed as the base classifier, and RSS and Dagging Conflict of interest The authors declare that they have no known com-
were used as the meta-classifiers among the ensemble mod- peting monetary interests or personal relationships that could have in-
els. This work has much opportunity to be expanded upon fluenced the work reported in this paper.
in the future. In the future, new variables and indices can be
added to provide a more accurate representation of drought Consent for Publication The authors authorized the publication of this
manuscript.
vulnerability. Deep learning techniques are being employed
in a variety of sectors nowadays. The creation of drought Consent To Participate .
vulnerability maps may, in the future, be accomplished by The authors declare that this research will be used only for scientific
deep learning techniques. Academics can stay up to date purposes and will not be passed on to third parties.
with the latest developments in drought prediction systems
and provide their insights to enhance them further (Madri-
gal et al. 2018). References
Achite M, Elshaboury N, Jehanzaib M, Vishwakarma DK, Pham QB,
Anh DT, Abdelkader EM, Elbeltagi A (2023) Performance of
machine learning techniques for meteorological drought forecast-
ing in the Wadi Mina basin. Algeria Water 15(4):765
Antronico L, De Pascale F, Coscarelli R, Gullà G (2020) Landslide
risk perception, social vulnerability, and community resilience:

13
A Data-Driven Approach to Spatial Drought Mapping Using Machine Learning

the case study of Maierato (calabria, Southern Italy. Int J Disaster Mann HB (1945) Nonparametric tests against trend. Econometrica: J
Risk Reduct 46:101529 Econometric Soc, pages 245–259
Antwi-Agyei P, Fraser ED, Dougill AJ, Stringer LC, Simelton E Markuna S, Kumar P, Ali R, Vishwkarma DK, Kushwaha KS, Kumar
(2012) Mapping the vulnerability of crop production to drought R, Singh VK, Chaudhary S, Kuriqi A (2023) Application of inno-
in Ghana using rainfall, yield, and socio-economic data. Appl vative machine learning techniques for long-term rainfall predic-
Geogr 32(2):324–334 tion. Pure Appl Geophys 180(1):335–363
Arabameri A, Asadi Nalivan O, Pal C, Chakrabortty S, Saha R, Lee McKee TB, Doesken NJ, Kleist J et al (1993) The relationship of
A, Pradhan S, B., and, Bui T, D (2020) Novel machine learning drought frequency and duration to time scales. In Proceedings
approaches for modeling the gully erosion susceptibility. Remote of the 8th Conference on Applied Climatology, volume 17, pages
Sens 12(17):2833 179–183. California
Ashraf M, Routray JK (2015) Spatio-temporal characteristics of pre- Mehr AD, Vaheddoost B, Mohammadi B (2020) Enn-sa: A novel
cipitation and drought in Balochistan Province. Pakistan Nat Haz- neuro-annealing model for multi-station drought prediction.
ards 77:229–254 Comput Geosci 145:104622
Aziz MA, Hossain AZ, Moniruzzaman M, Ahmed R, Zahan T, Azim Meshram SG, Hasan MA, Nouraki A, Alavi M, Albaji M, Meshram C
S, Qayum MA, Mamun A, Kader MA, M. A., and, Rahman (2023) Machine learning prediction of sediment yield index. Soft
NMF (2022) Mapping of agricultural drought in Bangladesh Comput 27(21):16111–16124
using geographic information system (gis). Earth Syst Environ Moayedi H, Mehrabi M, Kalantar B, Abdullahi Mu’azu M, Rashid
6(3):657–667 A, Foong AS, L. K., and, Nguyen H (2019) Novel hybrids of
Babaei H, Araghinejad S, Hoorfar A (2013) Developing a new method adaptive neuro-fuzzy inference system (anfis) with several
for Spatial assessment of drought vulnerability (case study: Z metaheuristic algorithms for Spatial susceptibility assessment
ayandeh-r Ood river basin in Iran). Water Environ J 27(1):50–57 of seismic-induced landslide. Geomatics Nat Hazards Risk
Barzegar R, Razzagh S, Quilty J, Adamowski J, Pour HK, Booij MJ 10(1):1879–1911
(2021) Improving galdit-based groundwater vulnerability predic- Mosavi A, Shirzadi A, Choubin B, Taromideh F, Hosseini FS, Borji M,
tive mapping using coupled resampling algorithms and machine Shahabi H, Salvati A, Dineva AA (2020) Towards an ensemble
learning models. J Hydrol 598:126370 machine learning model of random subspace-based functional
Cheng L, Hoerling M, AghaKouchak A, Livneh B, Quan X-W, Eis- tree classifier for snow avalanche susceptibility mapping. IEEE
cheid J (2016) How has human-induced climate change affected Access 8:145968–145983
California’s drought. risk? J Clim 29(1):111–120 Nabipour N, Dehghani M, Mosavi A, Shamshirband S (2020) Short-
Dikshit A, Pradhan B, Alamri AM (2021) Pathways and challenges of term hydrological drought forecasting based on different nature-
the application of artificial intelligence to geohazards modeling. inspired optimization algorithms hybridized with artificial neural
Gondwana Res 100:290–301 networks. IEEe Access 8:15210–15222
Dong X, Yu Z, Cao W, Shi Y, Ma Q (2020) A survey on ensemble Nguyen V-T, Tran TH, Ha NA, Ngo VL, Nadhir A-A, Tran VP, Nguyen
learning. Front Comput Sci 14:241–258 D, Amini HMAM, Prakash A, I., et al (2019) Gis based novel
Elbeltagi A, Kumar M, Kushwaha NL, Pande CB, Ditthakit P, Vish- hybrid computational intelligence models for mapping land-
wakarma DK, Subeesh A (2023) Drought indicator analysis and slide susceptibility: a case study at Da Lat City. Vietnam Sustain
forecasting using data driven models: case study in Jaisalmer, 11(24):7118
India. Stoch Env Res Risk Assess 37(1):113–131 Nhu V-H, Janizadeh S, Avand M, Chen W, Farzin M, Omidvar E,
Gökkaya H, Kellegöz T (2017) Ahp, Topsis and Hungarian algorithm Shirzadi A, Shahabi H, Clague J, Jaafari J, A., et al (2020a) Gis-
based decision support model for staff appointment. Endüstri based gully erosion susceptibility mapping: a comparison of com-
Mühendisliği Dergisi 28(1):2–18 putational ensemble data mining models. Appl Sci 10(6):2039
Hadisuwito A, Hassan F (2018) Selection drought index calculation Nhu V-H, Shahabi H, Nohani E, Shirzadi A, Al-Ansari N, Bahrami S,
methods using electre, Topsis, and analytic hierarchy process. Int Miraki S, Geertsema M, Nguyen H (2020b) Daily water level pre-
J Eng Technol 7(436):1413–1418 diction of Zrebar lake (Iran): a comparison between m5p, random
Hoque M, Pradhan B, Ahmed N, Alamri A (2021) Drought vulner- forest, random tree, and reduced error pruning trees algorithms.
ability assessment using Geospatial techniques in Southern ISPRS Int J Geo-Information 9(8):479
Queensland. Australia Sens 21(20):6896 Onyari EK, Ilunga F (2013) Application of Mlp neural network and
Hwang C-L, Yoon K, Hwang C-L, Yoon K (1981) Methods for multi- m5p model tree in predicting streamflow: A case study of Luvu-
ple attribute decision making. Multiple attribute decision making: vhu catchment, South Africa. Int J Innov Manage Technol 4(1):11
methods and applications a state-of-the-art survey, pages 58–191 Palchaudhuri M, Biswas S (2016) Application of Ahp with Gis in
Khan F (2003) Geography of Pakistan: population, economy and drought risk assessment for puruliya district, India. Nat Hazards
environment 84:1905–1920
Liang L, ZHAO S-h, Chong QINZ-hHEK-x, LUO C, Y.-x., and Pham BT, Prakash I, Bui DT (2018) Spatial prediction of landslides
ZHOU, X.-d (2014) Drought change trend using modis Tvdi and using a hybrid machine learning approach based on random sub-
its relationship with climate factors in China from 2001 to 2010. space and classification and regression trees. Geomorphology
J Integr Agric 13(7):1501–1508 303:256–270
Lin J, Jiang Y, Zheng W, Wang H, Shi Y et al (2008) Automatic estab- Quinlan JR (1992) Learning with Continuous Classes. Proceedings of
lishment of the initial black start schemes for power systems. Australian Joint Conference on Artificial Intelligence. Open Jour-
Autom Electr Power Syst 32(2):72–75 nal of Geology, Hobart 16–18 November 1992, 343–348
Madrigal J, Solera A, Suárez-Almiñana S, Paredes-Arquiola J, Andreu Roodposhti MS, Safarrad T, Shahabi H (2017) Drought sensitivity
J, Sanchez-Quispe ST (2018) Skill assessment of a seasonal fore- mapping using two one-class support vector machine algorithms.
cast model to predict drought events for water resource systems. Atmos Res 193:73–82
J Hydrol 564:574–587 Roy J, Saha S (2021) Integration of artificial intelligence with meta-
Malik A, Kumar A, Salih SQ, Kim S, Kim NW, Yaseen ZM, Singh classifiers for the gully erosion susceptibility assessment in Hin-
VP (2020) Drought index prediction using advanced fuzzy logic glo river basin, Eastern India. Adv Space Res 67(1):316–333
model: regional case study over Kumaon in India. PLoS ONE Saha S, Kundu B, Paul GC, Mukherjee K, Pradhan B, Dikshit A, Abdul
15(5):e0233280 Maulud KN, Alamri AM (2021) Spatial assessment of drought

13
F. Fatima et al.

vulnerability using fuzzy-analytical hierarchical process: a case Wu J, Liu Z, Yao H, Chen X, Chen X, Zheng Y, He Y (2018) Impacts
study at the Indian state of Odisha. Geomatics. Nat Hazards Risk of reservoir operations on multi-scale correlations between
12(1):123–153 hydrological drought and meteorological drought. J Hydrol
Sammut C, Webb G (2017) Random subspace method. Encyclopedia 563:726–736
of machine learning and data mining. Springer US, Boston, pp Yevjevich VM et al (1967) An objective approach to definitions and
1055–1055 investigations of continental hydrologic droughts, vol 23. Colo-
Shaw R (2015) Floods in the Hindu Kush region: causes and socio- rado State University Fort Collins, CO, USA
economic aspects. Mountain hazards and disaster risk reduction, Zahraie B, Nasseri M (2011) Basin scale meteorological drought fore-
pages 33–52 casting using support vector machine (svm). In International con-
Shen H, Chen Y, Wang Y, Xing X, Ma X (2020) Evaluation of the ference on drought management strategies in arid and semi arid
potential effects of drought on summer maize yield in the Western regions. Muscat, Oman, pages 1–16
Guanzhong plain. China Agron 10(8):1095 Zhang Y, Liu X, Jiao W, Zeng X, Xing X, Zhang L, Yan J, Hong Y
Svetnik V, Liaw A, Tong C, Culberson JC, Sheridan RP, Feuston BP (2021) Drought monitoring based on a new combined remote
(2003) Random forest: a classification and regression tool for sensing index across the transitional area between humid and arid
compound classification and Qsar modeling. J Chem Inf Comput regions in China. Atmos Res 264:105850
Sci 43(6):1947–1958 Zounemat-Kermani M, Batelaan O, Fadaee M, Hinkelmann R (2021)
Talukdar S, Ghose B, Shahfahad, Salam R, Mahato S, Pham QB, Linh Ensemble machine learning paradigms in hydrology: A review. J
NTT, Costache R, Avand M (2020) Flood susceptibility modeling Hydrol 598:126266
in Teesta river basin, Bangladesh using novel ensembles of bag- Zuo D, Cai S, Xu Z, Peng D, Kan G, Sun W, Pang B, Yang H (2019)
ging algorithms. Stoch Env Res Risk Assess 34:2277–2300 Assessment of meteorological and agricultural droughts using in-
Tang H, Wen T, Shi P, Qu S, Zhao L, Li Q (2021) Analysis of charac- situ observations and remote sensing data. Agric Water Manage
teristics of hydrological and meteorological drought evolution in 222:125–138
Southwest China. Water 13(13):1846
Tian L, Yuan S, Quiring SM (2018) Evaluation of six indices for Publisher’s Note Springer Nature remains neutral with regard to juris-
monitoring agricultural drought in the south-central united States. dictional claims in published maps and institutional affiliations.
Agric for Meteorol 249:107–119
Ting KM, Witten IH (1997) Stacking bagged and dagged models Springer Nature or its licensor (e.g. a society or other partner) holds
Topcu E (2022) Drought analysis using the entropy weight-based Top- exclusive rights to this article under a publishing agreement with the
sis method: A case study of Kars, Turkey. Russ Meteorol Hydrol author(s) or other rightsholder(s); author self-archiving of the accepted
47(3):224–231 manuscript version of this article is solely governed by the terms of
Trajković S, Gocić M, Misic D, Milanovic M (2020) Spatio-temporal such publishing agreement and applicable law.
distribution of hydrological and meteorological droughts in the
South Morava basin. Nat Risk Manage Engineering: NatRisk
Project, pages 225–242
Wable PS, Jha MK, Shekhar A (2019) Comparison of drought indi-
ces in a semi-arid river basin of India. Water Resour Manage
33:75–102

13

You might also like