0% found this document useful (0 votes)
4 views13 pages

Rainfall Forecasting Model Using Machine Learning

Uploaded by

Wubishet Hailu
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views13 pages

Rainfall Forecasting Model Using Machine Learning

Uploaded by

Wubishet Hailu
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Ain Shams Engineering Journal xxx (xxxx) xxx

Contents lists available at ScienceDirect

Ain Shams Engineering Journal


journal homepage: [Link]

Electrical Engineering

Rainfall forecasting model using machine learning methods: Case study


Terengganu, Malaysia
Wanie M. Ridwan a,b, Michelle Sapitang a,b, Awatif Aziz a, Khairul Faizal Kushiar c, Ali Najah Ahmed d,⇑,
Ahmed El-Shafie e,f
a
UNITEN R&D Sdn. Bhd., Universiti Tenaga Nasional (UNITEN), Kajang 43000, Selangor Darul Ehsan, Malaysia
b
Department of Civil Engineering, College of Engineering, Universiti Tenaga Nasional (UNITEN), 43000 Kajang, Selangor, Malaysia
c
Asset Management Department, Generation Division, Tenaga Nasional Berhad, 59200 Kuala Lumpur, Malaysia
d
Institute for Energy Infrastructure (IEI), Universiti Tenaga Nasional (UNITEN), Kajang 43000, Selangor Darul Ehsan, Malaysia
e
Department of Civil Engineering, Faculty of Engineering, University of Malaya (UM), 50603 Kuala Lumpur, Malaysia
f
National Water Center (NWC), United Arab Emirates University, Al Ain P.O. Box. 15551, UAE

a r t i c l e i n f o a b s t r a c t

Article history: Rainfall plays a main role in managing the water level in the reservoir. The unpredictable amount of rain-
Received 1 July 2020 fall due to the climate change can cause either overflow or dry in the reservoir. In this study, several mod-
Revised 7 September 2020 els and methods were applied to predict the rainfall data in Tasik Kenyir, Terengganu. The comparative
Accepted 9 September 2020
study was conducted focusing on developing and comparing several Machine Learning (ML) models, eval-
Available online xxxx
uating different scenarios and time horizon, and forecasting rainfall using two types of methods. Data
involved for this research consist of taking the average rainfall from 10 stations around the study area
Keywords:
using Thiessen polygon to weight the station area and projected rainfall. The forecasting model uses four
Forecasting rainfall
Machine learning algorithms
different ML algorithms, which are Bayesian Linear Regression (BLR), Boosted Decision Tree Regression
Boosted decision tree regression (BDTR), Decision Forest Regression (DFR) and Neural Network Regression (NNR). On the other hand,
Decision forest regression the rainfall was predicted on different time horizon by using different ML’s algorithms which is method
Neural network regression 1 (M1): Forecasting Rainfall Using Autocorrelation Function (ACF) and method 2 (M2): Forecasting
Bayesian linear regression Rainfall Using Projected Error. In M1, the best regression developed for ACF is BDTR since it has the high-
est coefficient of determination, R2, after tuning the hyperparameter. The results show coefficient
between 0.5 and 0.9 with the highest of each scenarios for daily (0.9739693), weekly (0.989461), 10-
days (0.9894429) and monthly (0.9998085). In M2, overall model performances show that normalization
using LogNormal is preferably giving a good result of each categories except for 10-days with BDTR and
DFR are the most acceptable result than NNR and BLR. It is concluded that, two different methods have
been applied with different scenarios and different time horizons, and M1 shows a rather high accuracy
than M2 using BDTR modeling.
Ó 2020 THE AUTHORS. Published by Elsevier BV on behalf of Faculty of Engineering, Ain Shams Uni-
versity. This is an open access article under the CC BY-NC-ND license ([Link]
by-nc-nd/4.0/).

1. Introduction

1.1. Background
⇑ Corresponding author.
E-mail addresses: [Link]@[Link] (W.M. Ridwan), michelle. Malaysia is described as a country that has hot climate all
sapitang@[Link] (M. Sapitang), [Link]@[Link] (A. Aziz), mah- through of the year since it is placed close to the equator [1]. Cli-
foodh@[Link] (A.N. Ahmed), elshafie@[Link] (A. El-Shafie). mate change causes parts of the water cycle accelerate as global
Peer review under responsibility of Ain Shams University. warming temperatures raise the rate of evaporation around the
world. Evaporation, also called evapotranspiration, is defined as
the combination of evaporation and plant transpiration from the
land or ground to the atmosphere. The factors that affect the evap-
Production and hosting by Elsevier otranspiration are air temperature, wind speed, vapor pressure,

[Link]
2090-4479/Ó 2020 THE AUTHORS. Published by Elsevier BV on behalf of Faculty of Engineering, Ain Shams University.
This is an open access article under the CC BY-NC-ND license ([Link]

Please cite this article as: W.M. Ridwan, M. Sapitang, A. Aziz et al., Rainfall forecasting model using machine learning methods: Case study Terengganu,
Malaysia, Ain Shams Engineering Journal, [Link]
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

Notation

M1 Method 1 ENN-KHA Elman Neural Network with Hybrid of Krill-Herd Algo-


M2 Method 2 rithm
ML Machine Learning ENN Elman Neural Network
ECR East Coast Region SVR-FA Support Vector Regression with Hybrid of Firefly Algo-
GIS Geographical Information System rithm
PCA Principal Component Analysis ENN-FA Elman Neural Network with Hybrid of Firefly Algorithm
RF Random Forest MLP-ANN Multilayer Perceptron Neural Network
BLR Bayesian Linear Regression MAE Mean Absolute Error
BDTR Boosted Decision Tree Regression RMSE Root Mean Square Error
DFR Decision Forest Regression RSE Relative Squared Error
NNR Neural Network Regression RAE Relative Absolute Error
MLP Multiple Layer Perceptron R Coefficient of Determination / Correlation Coefficient
SMFM Single Mode Forecasting Model ACF Autocorrelation Function
MMFM Multiple Mode Forecasting Model PACF Partial Autocorrelation Function
ANFIS Neuron Fuzzy Inference System mm millimeter
SFLA Shuffled Frog-Leaping Algorithm ha hectare
SVR Support Vector Regression m meter
SVR-KHA Support Vector Regression with Hybrid of Krill-Herd
Algorithm

relative humidity, soil moisture, type of crop and crop growth sea- The monsoon occurs at the East Coast of Peninsular Malaysia,
son. On average, more evaporation leads to high precipitation rates which is Kelantan, Terengganu and Pahang, but this study focus
and the impacts can be seen in many parts of Malaysia, but it is not on the surrounding area of Tasik Kenyir, Hulu Terengganu. In con-
evenly distributed. Some areas may experience more massive than junction with describing the weather and climate of Tasik Kenyir,
average precipitation and the other regions may become prone to the nearest meteorological station was used to acquire data. The
droughts as the traditional locations of monsoon occurs are in parameters considered were rainfall amount in millimeters (mm)
the eastern part. The weather and climate are commonly described with daily data covers the period from year 1985 to year 2019. Late
through a few meteorological parameters and the foremost imper- December 2014, Terengganu experienced its most extreme rainfall
ative meteorological parameter is rainfall intensity or rainfall events which attributed to climate change [8]. Therefore, the pre-
amount [2]. Exuberant rain usually causing surges and avalanches, sent of this study is to predict the amount of rainfall for future to
also leads to natural disaster. One of the major focuses of climate prevent and help in flood forecasting.
change study is to understand whether there is a change in the The East Coast Region (ECR) experienced a few monsoon sea-
occurrence frequency and strength of heavy rainfall events. More- sons which known as southwest monsoon (dry seasons), northeast
over, rainfall form main input to the river basin where it affected monsoon (wet seasons) and inter monsoon [9]. The southwest
the water capacity and release of a stream particularly during monsoon occurs from May to August, northeast monsoon occurs
the torrential rainfall event. from November to February and inter monsoon occurs from
The rainfall-runoff relationship is one of the most complex September to October and March to April [10]. Climate of Tereng-
hydrological phenomena due to existence of complex non-linear ganu is best described by different monsoon seasons. The data
connection within the change of precipitation into runoff. This management was performed in dividing the data into five charac-
action is very challenging to assimilate, owing to the existence of teristic which is daily, weekly, 10-days, and monthly.
extensive number of factors included in the trial of physical pro-
cess [4]. Subsequently, its definite modelling is imperative for 1.2. Literature review
water resource improvement and management and forecast of cli-
mate change like dry season and surges [5]. Based on the involve- In forecasting rainfall, there are two dominant approaches
ment of physical angles, rainfall—runoff models are classified as which is conceptual modelling and system theoretical modelling
either system theoretic model or physical-based models. The phys- [11,12]. Conceptual modelling is widely applied in hydrological
ical -based models require the significant data almost the system forecasting because it considered to estimate within the physical
mechanism as well as its parameters but different with system mechanism which oversee the hydrologic process and ordinarily
theoretical models which do not concern much about the physical based on the features and expertise of a catchment. Unfortunately,
processes of the problem because these models based on the rain- this approach may not be attainable for rainfall forecasting due to
fall and runoff data. Its look for characterize nonlinearity and non- critical calibration data to reasoning rainfall is not effectively to be
stationary conduct from those information by the utilize of trans- collected and the work of rainfall calculations need advanced
mission functions [6]. Based on [7], Artificial Neural Network numerical tool [13]. System theoretical methods are applied map-
(ANN) have received global attention for rainfall-runoff modelling ping models to define the relationship between inputs and outputs
because of their capability to capture high degree of non-linearity while neglecting the physical structure processes. The most popu-
and climate change of relationship between the hydrological vari- lar approaches is ARMAX by [14] in time series forecasting how-
ables without fully understanding the processes beneath. Mean- ever this model incapable to predict changes that are not based
while, this study makes a difference to constructed and on the past data, especially for the nonlinearity of rainfall corre-
compared four new different ML techniques which is Neural Net- lated variables.
work Regression (NNR), Boosted Decision Tree Regression (BDTR), Recently, because of the numerous progress within the fields of
Bayesian Linear Regression (BLR), and Decision Forest Regression pattern recognition methodology, there is various tools to forecast
(DFR) in forecasting rainfall. rainfall easily rather than previously used conventional approach
2
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

of linear mathematical relationships supported operator experi- capable in predicting rainfall with acceptable level of precision
ence, mathematical curves and guidelines [15]. Machine Learning and in reducing the error in the dataset of the projected rainfall
tool is widely applied to unravel hydrological problems including from climate change model with the expected observable rainfall.
rainfall forecasting. The important of this modelling is that the skill
of the software to plot the input-output patterns without afore- 1.4. Objectives
mentioned expertise of the factors affecting the forecast parame-
ters [16,17]. Subsequently, researchers have since started to hone The major aim of the study is to develop and compare several
this ML to predict form of the modelling approaches and parameter ML models, namely Boosted Decision Tree Regression (BDTR),
to improve precision and unwavering quality of depict the predict- Decision Forest Regression (DFR), Bayesian Linear Regression
ing model. In addition, researcher have used ML for various param- (BLR) and Neural Network Regression (NNR) for the prediction of
eter such as predicting daily streamflow using multi-layer rainfall in Kenyir Lake, Hulu Terengganu. Secondly, this study also
perceptron (MLP) [18] and neuro-fuzzy inference system (ANFIS) assesses and evaluate different scenarios and time horizons to find
with shuffled frog-leaping algorithms (SFLA) [19], predicting daily the most accurate and reliable model and the effectiveness of these
evapotranspiration using support vector regression (SVR) [20], algorithms in learning the sole input of rainfall pattern. Thirdly, the
estimating of solar radiation using support vector regression study also uses two method to predict the rainfall which are fore-
(SVR) with hybrid of krill-herd algorithm (SVR-KHA) [21] and esti- casting rainfall using Autocorrelation Function (ACF) and forecast-
mating soil temperature using support vector regression (SVR) and ing rainfall using projected error. This study is a vital to enhance
elman neural network (ENN) with hybrid of firefly algorithm (SVR- the water management, especially in Kenyir Lake area.
FA, ENN-FA) and krill herd algorithm (SVR-KHA, ENN-KHA) [22]
and predicting sea-level using multilayer perceptron neural net-
work (MLP-ANN) and ANFIS [23]. 2. Materials and methods
Researcher [24–26] are used ANN to forecast rainfall with each
different method where [24] construct rainfall simulation model 2.1. Case study and data description
and provide accurate rainfall forecasting information, temporal
and spatial distribution, [25] used short-term rainfall for urban Data for this study consist of 10 stations covering the Kenyir
catchment and the result show that the ANN model with lower Lake area in Hulu Terengganu, Terengganu as shown in Fig. 1.
lag outflanked in terms of forecasting exact index, [26] forecast The basic statistical parameters are presented in Table 1. Table 1
daily rainfall with resilient propagation learning algorithm and shows that the standard deviation for each station varies between
the numerical results shown that proposed model is predominant 18 mm and 30 mm. All the stations have recorded rainfall of 0 mm
to a numerous regression model in terms of forecasting precision as the minimum and the maximum rainfall is 539.5 mm in Station
indices. [27] forecast rainfall using seasonal ARIMA by arrange to 7, followed by Station 1 (455.5 mm) and Station 2 (440 mm). The
monthly rainfall data and it turns out ARIMA (0,0,1)(1,1,1) was to maximum rainfall range for all the station in between the range
be the most effective to predict future precipitation with a 95% of 325.5 mm to 539.5 mm.
confidence interval. [28] aim to compare random forest (RF) and For the 10 stations, each station should have 3,455 daily rainfall
support vector machine (SVM) for real time radar derived rainfall values. But most of the rainfall station values (except for two sta-
forecasting with two different forecast models which is single tions; Station 1 and Station 7) are not indicated due to sensor or
mode forecasting model (SMFM) and multiple mode forecasting technical issue. The most common approach is to delete individu-
model (MMFM). The result show that SVN based SMFM is more als and/or variables containing missing observations. However,
effective than RF based SMFM because in most cases RF based this loss of information reduces the ability to detect patterns
SMFM underestimates the observed radar derived rainfalls. and introduces biases if data are not missing completely at ran-
Meanwhile, researcher [29] use Emotional Neural Network dom [32]. Principal Component Analysis commonly simplifies as
(ENN) and Artificial Neural Network (ANN) for modelling rainfall- PCA, is the oldest and widely used machine learning technique. It
runoff in the Sone Command, Bihar as this area experiences flood explores correlations between variables and determines the com-
due to heavy rainfall. Then ENN model give the best result for bination of values that better represents variations in outcomes.
rainfall-runoff discharge with determination of correlation, Such integrated feature values are used to create a more compact
R2 = 0.879. Researcher [30] conclude that Bayesian Network (BN) feature space called the principal (key) components [33,34].
and ANN models can be successfully water quality modelling and In this study, the Thiessen Polygon method calculates the prox-
forecasting but the most great in dealing with time series is BN imity area around the point with other points. It also can help to
including incomplete or missing data. Therefore, this study con- calculate station rainfall weight and average rainfall based on each
structed and compared four new different ML techniques which station. However, this method is suitably used for flat areas [23].
is Boosted Decision Tree Regression (BDTR), Bayesian Linear The total area for each rainfall stations using the Thiessen Polygon
Regression (BLR), Neural Network Regression (NNR) and Decision method is shown in Table 2 below.
Forest Regression (DFR) in forecasting rainfall. P1 A1 þ P2 A2 þ P3 A3    : þ Pn An
P¼ ð1Þ
A1 þ A2 þ A3    þ An
1.3. Problem statement
where P 1 is the daily rainfall values for first station, A1 is the area of
Rainfall forms the primary input to the river basin, affecting the Thiessen polygon area for first station, Pn is the daily rainfall values
water capacity a stream, particularly during the torrential rainfall at nth station and An = the area of Thiessen polygon area at nth
event [3]. Moreover, one of the major focuses of climate change station.
study is to understand whether there is an extreme changes in The average rainfall for the 10 stations is then calculated using
the occurrence and frequency of heavy rainfall events. The accu- the Equation in (1) referred to Lal [35] to get the average of daily
racy level of the ML models used in predicting rainfall based on rainfall. The basic statistical parameters for 3,455 daily value of
historical data has been one of the most critical concerns in hydro- average rainfall from January 2010 until June 2019, namely, mean,
logical studies [31]. An accurate ML forecasting model could give median, standard deviation (S.D.), range, minimum and maximum
early alerts of severe weather to help prevent natural disasters of the average rainfall data, are presented in Table 3. For this study,
and destruction. Hence, there is needs to develop ML algorithms the data also use projected rainfall to predict the error. From
3
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

Fig. 1. Rainfall points within Kenyir Lake areas.

Table 1
Rainfall station with its’descriptive analysis.

Station No. Mean (mm) Standard Deviation (mm) Range (mm) Minimum (mm) Maximum (mm)
1 12.73864 30.26158 455.5 0 455.5
2 11.26588 28.4316 440 0 440
3 9.788515 23.91329 344.5 0 344.5
4 10.8276 26.01967 420 0 420
5 8.882498 22.47946 325.5 0 325.5
6 7.987888 18.84153 352.5 0 352.5
7 9.811867 26.31595 539.5 0 539.5
8 9.587695 23.36292 360 0 360
9 9.299299 21.94225 339.5 0 339.5
10 10.92149 28.89012 424 0 424

4
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

Table 2 the number of classes [42]. A neural network (NN) is defined by its
Polygon station area of for the Kenyir Lake structure, including the number of hidden layers, the number of
catchment.
nodes in each hidden layer, how the layers are linked, which acti-
Station No. Polygon Area (km2) vation function is used, and the weights. NN’s are widely known for
1 228.2137 use in deep learning and modeling complex problems such as
2 172.3874 image recognition. They are easily adapted to regression problems.
3 129.5302 Thus, NNR is suited to situations where a more traditional regres-
4 153.1546
5 173.3922
sion model cannot fit a solution.
6 413.7608 Bayesian Inference is used in the Bayesian approach, unlike lin-
7 211.0905 ear regression [43]. Prior parameter information is combined with
8 354.8176 a likelihood function to generate parameter estimates, which
9 150.7165
means the forecast distribution evaluates the likelihood of a value
10 331.7657
P y given x for a particular w, by means of likelihood by current
Area ¼2318.8292
belief about w given data (y, X). Finally, sum up all possible values
of w [43]. BLR enables the survival of insufficient data or incor-
rectly distributed data by a fairly natural process. The major ad-
vantage is that, by this Bayesian processing, you recover the
Table 3, there is a huge difference for the maximum rainfall, about
whole range of inferential solutions, rather than a point estimate
517 mm.
and a confidence interval as in classical regression.
To summarize, the choice of the proposed methods to imple-
2.2. Models used for forecasting ment rainfall forecasting model is the difficulty for mimicking
the rainfall process utilizing traditional model methods. This is
The forecasting model uses four different ML algorithms, which due to the fact that the rainfall behavior is affected by different
are Bayesian Linear Regression (BLR), Boosted Decision Tree stochastic and natural resources such as temperatures rise and
Regression (BDTR), Decision Forest Regression (DFR) and Neural the air becomes warmer, more moisture evaporates from land
Network Regression (NNR). and water to atmosphere and also climate change causes shifts
A BDTR is a classic method to create an ensemble of regression in air and ocean currents which change the weather pattern.
trees where each tree is dependent on prior tree [36]. In a simple
word, it is an ensemble learning method during which the errors
of the primary tree will be corrected by the second tree, the third
2.3. Optimization technique
tree corrects for the errors of the primary and second trees, then
so on. Predictions are based together on the whole set of trees
The optimization technique for M1 uses cross-validation with
which makes the prediction. The BDTR shows an exceptionally
tuning. This optimization technique provides the best precision
great in dealing with tabular data [37]. The advantages of BDTR
and helps to find problems with datasets by dividing data into
is it robust to missing data and normally allocates feature signifi-
some fold, then building and testing the models on each fold
cance scores. Usually BDTR perform better than DFR because it
[44]. Numbers of input also played a role because the more input
appears to be the chosen method with slightly better performance
being generated, the highest result of coefficient of determination.
than DFR in Kaggle competition [38]. Unlike DFR, BDTR may be
Meanwhile, the optimization technique for M2 uses varies
more prone to overfitting because the main purpose is to reduce
types of normalization. Normalization is a method frequently
bias and not variance. BDTR takes a longer time to develop model
applied as part of data preprocessing for ML. The purpose of nor-
since it has many hyperparameters to tune and trees are built
malization is to adjust the values of numeric columns in the data-
sequentially [39].
set to use a standard scale, without distorting discrepancies in the
A DFR is an ensemble of randomly trained decision trees [40]. It
spectrum of values or losing information [45]. It refers to scaling
works by constructing a vast number of decision trees at training
down the dataset so that the weight data lies between 0 and 1. This
time and produce an individual tree model of classes (classifica-
will help to easily compare corresponding normalized values from
tion) or mean forecast (regression) as the end of product. Each tree
varies datasets that can eliminate the effects of large and smaller
is developed using a random subset of features employing an irreg-
values of datasets. Basically, the normalization uses MinMax for-
ular subset of data that deviate the trees by appearing them
mula of z ¼ ½maxxminðxÞ
ðxÞminðxÞ
. In ML, this study has included several
diverse datasets. It has two parameters: the number of trees and
the number of selected features at each node. DFR good in generate types of normalization which are MinMax, LogNormal and ZScore
uneven data sets with missing variables since it is generally robust to optimize the modelling technique to get better result.
to overfitting. It also has lower classification errors and better f-
scores than decision trees, but it does not easily to interpret the
results [39]. Another disadvantage is the importance feature may 2.4. Scenarios
not be vigorous to variety within the preparing dataset.
NNR is a chain of linear operations scattered with various non- This study focuses on predicting the rainfall in different time
linear activation functions [41]. In general, the network has these horizons by using different ML algorithms. It has been divided into
defaults; the first layer is the input layer, the last layer is the out- two methods with different scenarios to find the robust model
put layer, and the hidden layer consisting of several nodes equal to mimicking the actual value. The methods are;

Table 3
Average and projected rainfall value with its’ descriptive analysis.

Description Mean (mm) Standard Deviation (mm) Minimum (mm) Maximum (mm) Median (mm) Total of Daily Data
Average Value of 10 Stations (Actual) 10.04186 20.80274 0 271.3711 3.57 3,455
Projected Rainfall 8.523774 24.14694 0 788.22 1.38 3,455

5
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

2.4.1. Method 1: Forecasting rainfall using Autocorrelation Function as output, where Rt-2 know and rainfall at Lag 2. Eq. (4) use Rt, Rt-1
(ACF) and Rt-2 as input and Rt-3 as output, where Rt-3 know and rainfall at
Autocorrelation Function (ACF) is imperative analytical tools Lag 3.
utilized with the time series analysis and forecasting [46]. Measur-
ing the statistical relationships between observations is the pri- (ii) Weekly ACF
mary purpose of these models in a single data series. ACF has a
big advantage of measuring the amount of linear dependence Rt ¼ Rt1 ð5Þ
between results of a time series which will be separated by a lag
k. By plot these results along with the confidence band, it will Rt ¼ Rt49 ð6Þ
define how well the present value of the series is related with its
past values. ACF deals with all these components like trends, sea- Rt þ Rt49 ¼ Rt50 ð7Þ
sonality, cyclic, and residual while finding correlations; hence it Refer to Fig. 2(b), there are three different scenarios (Eqa. (5)–
called a complete auto-correlation plot. It is used to guess the form (7)) is develop based on Lag 1, Lag 49 and Lag 50 represent as week
of the model and obtain approximate estimates of the parameters. 1, week 49 and week 50 respectively. This is because the autocor-
Fig. 2(a to e) shows five scenarios of historical rainfall pattern in relation function has more than 0.2. Therefore, Eq. (5) use Rt as
ACF. input and Rt-1 as output, where Rt is actual rainfall and Rt-1 is rain-
For daily, the analysis took historical rainfall data from year fall at Lag 1. Eq. (6) use Rt as input and Rt-49 as output, where Rt-49
2010 until year 2019. Fig. 2(a) shows data observed for 31 days know and rainfall at Lag 49. Eq. (7) use Rt and Rt-49 as input and Rt-
using January 2010 and shows that the first day of the month is
50 as output, where Rt-50 know and rainfall at Lag 50.
more correlate compare to other days. Fig. 2(b) only observe for
one year which is 2010 consisting of 52 weeks, so the graph only (iii) 10 Days ACF
shows 52 of lag where 1 lag represents one week. Example lag 1
represent the beginning of 7 days for 2010 and continuously till Rt ¼ Rt1 ð8Þ
lag 52 represents the last 7-days for 2010. Lag 1 to lag 4 show
result for January 2010 while lag 48 to lag 52 show result for Rt ¼ Rt34 ð9Þ
November 2010 and December 2010. Lag 1 to lag 4 and lag 52 give
high results above the upper line where the amount of rainfall is Rt þ Rt34 ¼ Rt35 ð10Þ
high compared to other weeks.
Meanwhile, 10-days correlogram graphs (Fig. 2(c)) only observe Rt þ Rt34 þ Rt36 ¼ Rt37 ð11Þ
for one year, year 2010, consisting 37 lags of 10 days. In the dia-
gram, each lag equivalent to 10 days. For example, lag 1 represents Refer to Fig. 2(c), there are four different scenarios is develop
the beginning of 10 days for year 2010 and continuously till lag 37 based on Lag 1, Lag 34, Lag 35 and Lag 37 shown in Eq. (8)–(11).
represent the last 10-days for year 2010. Lag 1 to lag 3 show result This is because the autocorrelation function has more than 0.2.
for January 2010 while lag 33 to lag 37 show result for November Therefore, Eq. (8) use Rt as input and Rt-1 as output, where Rt is
2010 and December 2010. This lag gives high results above the actual rainfall and Rt-1 is rainfall at Lag 1. Eq. (9) use Rt as input
upper line, where the amount of rainfall is high compared to other and Rt-34 as output, where Rt-34 know and rainfall at Lag 34. Eq.
lag. (10) use Rt and Rt-34 as input and Rt-35 as output, where Rt-35 know
Fig. 2(d) show monthly rainfall data for year 2010 and year and rainfall at Lag 35. Lastly Eq. (11) use Rt, Rt-34, Rt-35 and Rt-36 as
2011. There are 24 of lag represent 24 months where lag 1 to lag input and Rt-37 as output, where Rt-37 know and rainfall at Lag 37.
12 for 2011 and lag 13 to lag 24 for 2011. Lag 1 and 13 represent
January, lag 11 and 23 represent November, lag 12 and 24 repre- (iv) Monthly ACF
sent December. This correlogram shows that the data have sea-
Rt ¼ Rt1 ð12Þ
sonal dependencies and the same pattern over the years. For
example, December 2011 (Lag 24) dependent on December 2010
Rt ¼ Rt11 ð13Þ
(Lag 12), which means it is repeated. Based on the graph, there
are dependencies between November, December and January in
Rt þ Rt11 ¼ Rt12 ð14Þ
each year. Therefore, each year, the wet season occurs on Novem-
ber until January. But in some years, there are also dependencies
Rt þ Rt11 þ Rt12 ¼ Rt13 ð15Þ
between October and some years do not have dependencies in
October. Therefore, in some years, the wet season may occur start- Refer to Fig. 2(d), there are four different scenarios (Eq. (12)–
ing from October until January. (15)) is develop based on Lag 1, Lag 11, Lag 12 and Lag 13 represent
as January 2010, November 2010, December 2010 and January
2011 respectively. This is because the autocorrelation function
[Link]. Input selection for method 1.
has more than 0.2. Therefore, Eq. (12) use Rt as input and Rt-11 as
(i) Daily ACF
output, where Rt is actual rainfall and Rt-1 is rainfall at Lag 1. Eq.
Rt ¼ Rt1 ð2Þ (13) use Rt as input and Rt-11 as output, where Rt-11 know and rain-
fall at Lag 11. Eq. (14) use Rt and Rt-11 as input and Rt-12 as output,
Rt þ Rt1 ¼ Rt2 ð3Þ where Rt-12 know and rainfall at Lag 12. Lastly Eq. (15) use Rt, Rt-11
and Rt-12 as input and Rt-13 as output, where Rt-13 know and rainfall
at Lag 13.
Rt þ Rt1 þ Rt2 ¼ Rt3 ð4Þ
Refer to Fig. 2(a), there are three different scenarios is develop 2.4.2. Method 2: Forecasting rainfall using projected error
based on Lag 1, Lag 2 and Lag 3 shown in Eq. (2)–(4). This is This study aims to predict the error of projected rainfall to cor-
because the autocorrelation function has more than 0.2. Therefore, rect the future projected error. Data for this study include the pro-
Eq. (2) use Rt as input and Rt-1 as output, where Rt is actual rainfall jected rainfall consist of daily data from year 2010 until year 2099,
and Rt-1 is rainfall at Lag 1. Eq. (3) use Rt and Rt-1 as input and Rt-2 which the projection point (in Fig. 3).
6
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

Fig. 2. AutoCorrelation Function (ACF) for; (a) Daily; (b) Weekly; (c) 10-days; and (d) Monthly.

The main input idea for forecasting rainfall using projected which the error is mainly from subtracted the projected rainfall Rp
error is in Eq. (16). Where,Rp is projected rainfall as input for the with actual rainfall Ra . The rainfall error data divided into several
model. As output for the model, Ep is the projected rainfall error. scenarios; daily, weekly, 10-days, and monthly.
The main idea for this equation is to correct the projected rainfall
from year 2020 to 2099 after predicting the projected rainfall error Rp ¼ Ep ð16Þ

7
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

Fig. 3. Projected rainfall points.

2.5. Performance indicators 3. Results and discussion

In this study, different model performance indicators were used 3.1. Method 1: Forecating rainfall using Autocorrelation Function
as shown in Table 4 to signify the successful of scoring (datasets) (ACF)
has been by a trained model to mimicking the real values of the
output parameters. Table 5 shows the best models results to predict rainfall based
In a nutshell, the forecasting performance is better when the on ACF. In this study, there are four different regression has been
value of R2 is close to 1 and differing for RMSE and MAE used, which is Bayesian Linear Regression (BLR), Boosted Decision
because the model’s performance will be better if the value is Tree Regression (BDTR), Decision Forest Regression (DFR) and Neu-
close to 0 [50]. Fig. 4 showing the flow chart of the research ral Network Regression (NNR). All regression generates different
methodology. scenario for rainfall data divided into daily, weekly, 10-days and

8
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

Table 4 the accuracy of the proposed model can be seen after adding cross
Performance indicators to evaluate the model. validation and hyper-parameters tuning techniques. By adding this
Equation Performance indicator step into the ML, the proposed BDTR model gives the best accuracy
PN in predicting the rainfall where the value of coefficient of determi-
MAE ¼ Mean absolute error, MAE [47] re-
i¼1 jMSLp  MSLo j
1
n
flects the degree of absolute error nation in predicting daily rainfall it ranges between (0.5525075-
between the actual and forecasted 0.9739693), and for weekly rainfall prediction it ranges between
data. (0.8400668-0.989461), and for 10 days rainfall prediction it ranges
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
rP
n 2 Root Mean Square Error, RMSE [47]
RMSE ¼ i¼1
ðMSLp MSLo Þ
is compared between the actual and
between (0.8038288-0.9894429) and finally for monthly rainfall
N
forecasted data. prediction it ranges between (0.9174191-0.9998085). Based on
Pn
jp ai j Relative absolute error, RAE [48] is these results, it can be concluded that BDTR capable of predicting
RAE ¼ Pni¼1 i
i¼1
ja ai j the relative absolute difference be- the rainfall in different time horizon with acceptable level of accu-
tween actual and forecasted values.
Pn racy and the accuracy of the model improved when more inputs
ðp ai Þ2 Relative squared error, RSE [48]
RSE ¼ Pni¼1 i 2 included.
i¼1
ða ai Þ similarly normalizes the entire
squared error of the forecasted val-
ues.
Pn    
3.2. Method 2: Forecasting rainfall using projected error
MSLo MSLo ðMSLp MSLp Þ Coefficient of determination, R2 [49]
R2 ¼ qffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
P
i¼1
P
n  2 n 2 is showing the performance of fore-
i¼1
ðMSLo MSLo Þ i¼1
ðMSLp MSLp Þ
casting model where zero means the Table 6 summarizes the metrics for each top three best models
model is random while 1 means
there is a perfect fit.
for different scenarios using different normalization and data par-
titioning. In this context, different normalization such as ZScore,
LogNormal, and MinMax and data partitioning (80% and 90%) are
monthly. Hence, the best regression developed for ACF is BDTR investigated to obtain the optimal model with high level of accu-
since it has the highest coefficient of determination, R2. racy. Overall model performances show that normalization using
It can be seen from the table that without tuning, the BDTR LogNormal gives a good result for each category except for 10-
model did not perform well, whereas, noticeable improvement in days prediction. Comparison between four models used, BDTR

Fig. 4. The methodology of current research.

9
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

Table 5 and DFR, are the most acceptable result than NNR and BLR. The
Result for the best model in M1 using ACF. outcome of the results shows the best model for daily error predic-
Scenario Regression Model Coefficient of tion with R equal to 0.737978 is BDTR and for weekly rainfall error
Determination prediction where R equal to 0.7921. While for monthly rainfall
(a) Daily error prediction, DFR outperformed other models where R equal
Rt = Rt-1 BDTR Without 0.2458173 to 0.7623. However, for 10-days rainfall error prediction, NNR
tuning model with ZScore normalization outperformed other models in
With tuning 0.5525075
Rt + Rt-1 = Rt-2 BDTR Without 0.1383447
predicting the value of 10-days with acceptable level of accuracy
tuning where R is equal to 0.61728. It can be concluded, acceptable level
With tuning 0.8468193 of accuracy could be achieved by reducing the error in the dataset
Rt + Rt-1 + Rt-2 = Rt-3 BDTR Without 0.2020856 of the projected rainfall with the expected observable rainfall by
tuning
using BDTR integrated with LogNormal and by partitioning the
With tuning 0.9739693
data to 90% for training and 10% for testing.
(b) Weekly
Finally, Fig. 5 demonstrates the actual error and the predicted
Rt = Rt-1 BDTR Without 0.0002462
tuning error. The figure shows how well the proposed model can resemble
With tuning 0.8400668 the actual error between the observed and projected rainfall during
Rt = Rt-49 BDTR Without 0.1179256 the testing phases. The red line demonstrates the predicted value
tuning from the proposed ML algorithm, while blue line demonstrates
With tuning 0.8825647
the observed value of actual error. It can be seen that the proposed
Rt + Rt-49 = Rt-50 BDTR Without 0.235989
tuning model has acceptable level of accuracy for all four different time
With tuning 0.989461 horizons, however, highest level of accuracy obtained for the
(c) 10 Days weekly error (Fig. 5(b).
Rt = Rt-1 BDTR Without 1.0041807
tuning
With tuning 0.8038288 4. Conclusion
Rt = Rt-34 BDTR Without 0.1182632
tuning
With tuning 0.8949389
This paper focuses on two methods; (1) Forecasting rainfall
Rt + Rt-34 = Rt-35 BDTR Without 0.4616042 using Autocorrelation Function (ACF) based on the historical rain-
tuning fall data and (2) Forecasting rainfall using Projected Error based on
With tuning 0.9607741 historical and projected rainfall data. Both methods using different
Rt + Rt-34 + Rt-35 + Rt- BDTR Without 0.1377565
algorithms such as BDTR, DFR, BLR and NNR to identify the optimal
36 = Rt-37 tuning
With tuning 0.9894429 prediction for rainfall and different time horizons. The results pre-
sented that for M1, the result gets better with cross-validation with
(d) Monthly
Rt = Rt-1 BDTR Without 0.1163886 BDTR and tuning its parameter. The more input included to the
tuning model; the more accurate the model can perform. The best regres-
With tuning 0.9174191 sion developed for ACF is BDTR since it has the highest coefficient
Rt = Rt-11 BDTR Without 0.0514856
of determination, R2 (daily: 0.5525075, 0.8468193, 0.9739693;
tuning
With tuning 0.6941756
weekly: 0.8400668, 0.8825647, 0.989461; 10 days: 0.8038288,
Rt + Rt-11 = Rt-12 BDTR Without 0.5693955 0.8949389, 0.9607741, 0.9894429; and monthly: 0.9174191,
tuning 0.6941756, 0.9939951, 0.9998085) meaning the better rainfall pre-
With tuning 0.9939951 diction for the future. For method 2, a variation of result when
Rt + Rt-11 + Rt-12 = Rt-13 BDTR Without 0.4366365
using different normalization techniques and shows using LogNor-
tuning
With tuning 0.9998085 mal normalization with BDTR and DFR gives the best model mim-
icking the actual projected error. The best-predicting scenarios is

Table 6
Summary of top best three models for each scenario.

Regression Normalization Training (%) MAE RMSE RAE RSE R


(a) Daily Error
BDTR LogNormal 90 0.105467 0.160547 0.391675 0.262022 0.737978
DFR LogNormal 90 0.106306 0.177975 0.394791 0.321997 0.678003
BDTR LogNormal 80 0.113675 0.178906 0.432401 0.332167 0.667833
(b) Weekly Error
BDTR LogNormal 90 0.064627 0.117037 0.314988 0.2079 0.7921
DFR LogNormal 90 0.058325 0.126458 0.284272 0.242716 0.757284
DFR LogNormal 80 0.078724 0.160686 0.348225 0.349252 0.650748
(c) 10-Days Error
NNR ZScore 80 0.389449 0.672152 0.716514 0.38272 0.61728
BLR ZScore 80 0.417034 0.674603 0.767265 0.385516 0.614484
DFR MinMax 80 0.031469 0.053312 0.801608 0.461544 0.538456
(d) Monthly Error
DFR LogNormal 80 0.094573 0.156451 0.379418 0.2377 0.7623
DFR LogNormal 90 0.151092 0.219738 0.47616 0.367674 0.632326
DFR MinMax 90 0.048869 0.069285 0.609473 0.372461 0.627539

10
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

Fig. 5. Comparison between actual error and predicted error for; (a) Daily; (b) Weekly; (c) 10-days; and (d) Monthly.

the weekly error with R of 0.7921 which is the closest to 1. This dencies on ACF show that rainfall has almost a similar pattern
means the BDTR model can best predict weekly average error to every year from November to January, and this shows a correlation
correct the weekly average projected rainfall. In conclusion, between the predicting input and output. The findings of the cur-
method 1 is the best prediction for the rainfall, mimicking the rent study showed that standalone machine learning algorithms
actual values with the highest coefficient closer to 1. The depen- capable to predict the rainfall with acceptable level of accuracy,
11
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

however, more accurate rainfall prediction might be achieved by Goldstein H, Molenberghs G, Scott D, Smith A, Tsay R, Weisberg S, editors. John
Wiley & Sons; 2016.
proposing hybrid machine learning algorithms and with the inclu-
[15] Tokar AS, Markus M. Precipitation-runoff modelling using artificial neural
sion of different climate change scenarios. networks and conceptual models. J Hydrol Eng 2000;5(April):156–61.
[16] Najah AA, El-Shafie A, Karim OA, Jaafar O. Integrated versus isolated scenario
for prediction dissolved oxygen at progression of water quality monitoring
Funding stations. Hydrol Earth Syst Sci Discuss 2011;8(3):6069–112.
[17] Hipni A, El-shafie A, Najah A, Karim OA, Hussain A, Mukhlisin M. Daily
forecasting of dam water levels: comparing a Support Vector Machine (SVM)
This research and the APC was funded by TNB Seed Fund (U-TG-
model with Adaptive Neuro Fuzzy Inference System (ANFIS). Water Resour
RD-19-01) managed by UNITEN R&D Sdn Bhd. Manag 2013;27(10):3803–23. doi: [Link]
4.
[18] Mohammadi B, Ahmadi F, Mehdizadeh S, Guan Y, Pham QB, Linh NTT, et al.
Declaration of Competing Interest Developing novel robust models to improve the accuracy of daily streamflow
modeling. Water Resour Manag 2020.
The authors declare that they have no known competing finan- [19] Mohammadi B, Linh NTT, Pham QB, Ahmed AN, Vojteková J, Guan Y, et al.
Adaptive neuro-fuzzy inference system coupled with shuffled frog leaping
cial interests or personal relationships that could have appeared
algorithm for predicting river streamflow time series. Hydrol Sci J 2020:1.
to influence the work reported in this paper. [20] Mohammadi B, Mehdizadeh S. Modeling daily reference evapotranspiration
via a novel approach based on support vector regression coupled with whale
optimization algorithm. Agric Water Manag 2019;2020(237):106145.
Acknowledgments [21] Mohammadi B, Aghashariatmadari Z. Estimation of solar radiation using
neighboring stations through hybrid support vector regression boosted by krill
The authors would like to thank Asset Management Depart- herd algorithm. Arab J Geosci 2020;13(10):1–16.
[22] Moazenzadeh R, Mohammadi B. Assessment of bio-inspired metaheuristic
ment, Generation Division, Tenaga Nasional Berhad, Malaysia for optimisation algorithms for estimating soil temperature. Geoderma
providing this study with the rainfall data and permission to con- 2018;2019(353):152–71.
duct the study and the authors would like to acknowledge the [23] Muslim TO, Ahmed AN, Malek MA, Afan HA, Ibrahim RK, El-shafie A, et al.
Investigating the influence of meteorological parameters on the accuracy of
financial support received from TNB Seed Fund (U-TG-RD-19-01)
sea-level prediction models in Sabah. Malaysia 2020.
managed by UNITEN R & D Sdn. Bhd. In addition to that, the [24] French MN, Krajewski WF, Cuykendall RR. Rainfall forecasting in space and
authors would like to thank Professor [Link] Mohd Sidek time using a neural network. J Hydrol 1992;137(1–4):1–31. doi: [Link]
for providing the projected rainfall data which used in this study. org/10.1016/0022-1694(92)90046-X.
[25] Luk KC, Ball JE, Sharma A. A study of optimal model lag and spatial inputs to
artificial neural network for rainfall forecasting. J Hydrol 2000;227(1–
References 4):56–65. doi: [Link]
[26] Valverde Ramírez MC, De Campos Velho HF, Ferreira NJ. Artificial neural
network technique for rainfall forecasting applied to the São Paulo Region. J
[1] Ismail M, Suroto A, Abdullah S, Nagarani N, Arumugam K, Kumaraguru VJ, et al.
Hydrol 2005;301(1–4):146–62. doi: [Link]
Response of Malaysian local rice cultivars induced by elevated ozone stress
jhydrol.2004.06.028.
EnvironmentAsia genotoxicity assessment of mercuric chloride in the Marine
[27] Bari SH, Rahman MT, Hussain MM, Ray S. Forecasting monthly precipitation in
Fish Therapon Jaruba.
Sylhet city using ARIMA model. Civil Environ. 2015;7:1.
[2] Abdullah S, Ismail M. The weather and climate of tropical Tasik Kenyir,
[28] Yu PS, Yang TC, Chen SY, Kuo CM, Tseng HW. Comparison of random
Terengganu. In: Greater Kenyir landscapes: social development and
forests and support vector machine for real-time radar-derived rainfall
environmental sustainability: from ridge to reef. Springer International
forecasting. J Hydrol 2017;552:92–104. doi: [Link]
Publishing: Cham; 2019. p. 3–8. doi:10.1007/978-3-319-92264-5_1.
jhydrol.2017.06.020.
[3] Khairul Amri Kamarudin M, Abd Wahab N, Juahir H, Mohd Firdaus Nik Wan N,
[29] Kumar S, Roshni T, Himayoun D. A comparison of Emotional Neural Network
Barzani Gasim M, et al. The potential impacts of anthropogenic and climate
(ENN) and Artificial Neural Network (ANN) approach for rainfall-runoff
changes factors on surface water ecosystem deterioration at Kenyir Lake,
modelling. Civ Eng J 2019;5(10):2120–30. doi: [Link]
Malaysia. Int J Eng Technol 2018; 7 (3.14): 67. doi:10.14419/ijet.
2019-03091398.
v7i3.14.16864.
[30] Jafari Nodoushan E. Monthly forecasting of water quality parameters within
[4] Chang TK, Talei A, Alaghmand S, Ooi MPL. Choice of rainfall inputs for event-
bayesian networks: a case study of Honolulu, Pacific Ocean. Civ Eng J 2018;4
based rainfall-runoff modeling in a catchment with multiple rainfall stations
(1):188. doi: [Link]
using data-driven techniques. J Hydrol 2017;545:100–8. doi: [Link]
[31] Sumi SM, Zaman MF, Hirose H. A rainfall forecasting method using machine
10.1016/[Link].2016.12.024.
learning models and its application to the fukuoka city case. Int J Appl Math
[5] Chang TK, Talei A, Quek C, Pauwels VRN. Rainfall-runoff modelling using a self-
Comput Sci 2012;22(4):841–54. doi: [Link]
reliant fuzzy inference network with flexible structure. J Hydrol
0062-1.
2018;564:1179–93. doi: [Link]
[32] Nakagawa S, Freckleton RP. Missing inaction: the dangers of ignoring missing
[6] Bartoletti N, Casagli F, Marsili-Libelli S, Nardi A, Palandri L. Data-driven
data. Trends Ecol Evol 2008;23(11):592–6.
rainfall/runoff modelling based on a neuro-fuzzy inference system. Environ
[33] Jollife IT, Cadima J. Principal component analysis: a review and recent
Model Softw 2018;106:35–47. doi: [Link]
developments. Philos Trans R Soc A Math Phys Eng Sci 2016;374(2065).
envsoft.2017.11.026.
[34] Wold S, Esbensen K, Geladi P. Principal component analysis. Chemom Intell
[7] Ghumman AR, Ghazaw YM, Sohail AR, Watanabe K. Runoff forecasting by
Lab Syst 1987;2(1–3):37–52.
artificial neural network and conventional model. Alexandria Eng J 2011;50
[35] Lal BB, Al-mashidani G. A Technique for the determination of areal average
(4):345–50. doi: [Link]
rainfall/Une Méthode Pour La Détermination de Précipitation Moyenne Pour
[8] Climate changes impacts towards sedimentation rate at Terengganu River,
Une Zone Donnée a technique for the determination of areal average rainfall.
Terengganu, Malaysia | J Fundam Appl Sci [Link]
Hydrol Sci Sci Hydrol 1978;23. doi: [Link]
jfas/article/view/168262 (accessed Jun 12, 2020).
02626667809491823.
[9] (PDF) Multiple Linear Regression (MLR) models for long term Pm10
[36] Boosted Decision Tree Regression – Azure Machine Learning Studio | Microsoft
concentration forecasting during different monsoon seasons [Link]
Docs [Link]
[Link]/publication/318777794_Multiple_Linear_Regression_
module-reference/boosted-decision-tree-regression (accessed Oct 23, 2019).
MLR_models_for_long_term_Pm10_concentration_forecasting_during_
[37] Boosted Trees Regression | Turi Machine Learning Platform User Guide
different_monsoon_seasons (accessed Mar 27, 2020).
[Link]
[10] Peredaran S; Dan A, Monsun S, Daya B, Perairan D, Timur P, et al. Simulation of
[Link] (accessed Oct 23, 2019).
Southwest Monsoon Current Circulation and Temperature in the East Coast of
[38] Pros and cons of classical supervised ML algorithms - rmartinshort https://
Peninsular Malaysia 2014; 43.
[Link]/2019/02/24/pros-and-cons-of-classical-
[11] Duan A, Publisher Q. A global optimization strategy for efficient and effective
supervised-ml-algorithms/ (accessed Feb 26, 2020).
calibration of hydrologic models. Item Type Text; Dissertation-Reproduction
[39] HackingNote [Link]
(Electronic). The University of Arizona. 1991.
algorithms-pros-and-cons (accessed Feb 26, 2020).
[12] Luk KC, Ball JE, Sharma A. An application of artificial neural networks for
[40] Criminisi A. Decision forests: a unified framework for classification, regression,
rainfall forecasting. Math Comput Model 2001;33(6–7):683–93. doi: https://
density estimation, manifold learning and semi-supervised learning. Found
[Link]/10.1016/S0895-7177(00)00272-7.
TrendsÒ Comput Graph Vis 2011;7(2–3):81–227. doi: [Link]
[13] Duan Q, Sorooshian S, Gupta VK. Optimal use of the SCE-UA global
0600000035.
optimization method for calibrating watershed models. J Hydrol 1994;158
[41] Tan LK, McLaughlin RA, Lim E, Abdul Aziz YF, Liew YM. Fully automated
(3–4):265–84. doi: [Link]
segmentation of the left ventricle in cine cardiac MRI using neural network
[14] Box G, Jenkins G, Reinsel G, Ljung G. Fifth edition time series analysis
regression. J Magn Reson Imaging 2017;48(1):140–52.
forecasting and control. In: Balding D, Cressie N, Fitzmaurice G, Givens G,

12
W.M. Ridwan, M. Sapitang, A. Aziz et al. Ain Shams Engineering Journal xxx (xxxx) xxx

[42] Damian DC. A critical review on artificial intelligence models in hydrological learning view project machine learning techniques for rainfall prediction: a
forecasting how reliable are artificial intelligence models. 2019, No. July. review.
[43] Catal C, Ece K, Arslan B, Akbulut A. Benchmarking of regression algorithms and [47] Hyndman RJ, Koehler AB. Another look at measures of forecast accuracy. Int J
time series analysis techniques for sales forecasting. Balk J Electr Comput Eng Forecast 2006;22(4):679–88.
2019;7. [48] Miller TB, Kane M. The precision of change scores under absolute and relative
[44] Tune Model Hyperparameters – ML Studio (classic) – Azure | Microsoft Docs interpretations. Appl Meas Educ 2001;14(4):307–27.
[Link] [49] Nagelkerke NJD. A note on a general definition of the coefficient of
reference/tune-model-hyperparameters (accessed Mar 31, 2020). determination. Biometrika 1991;78(3):691–2.
[45] Normalize Data: Module Reference – Azure Machine Learning | Microsoft Docs. [50] Cheng CT, Feng ZK, Niu WJ, Liao SL. Heuristic methods for reservoir monthly
[46] Parmar A, Mistree K, Sompura M. Machine learning techniques for rainfall inflow forecasting: a case study of Xinfengjiang reservoir in pearl river, China.
prediction: a review indian sign language recognition view project machine Water (Switzerland) 2015;7(8):4477–95.

13

You might also like