Hybrid Model for Solar Radiation Prediction
Hybrid Model for Solar Radiation Prediction
A novel hybrid model for hourly global solar radiation prediction using
random forests technique and firefly algorithm
Ibrahim Anwar Ibrahim a,⇑, Tamer Khatib b
a
Department of Electrical, Electronic and Systems Engineering, Universiti Kebangsaan Malaysia, 43600 Bangi, Selangor, Malaysia
b
Energy Engineering and Environment Department, An-Najah National University, 97300 Nablus, Palestine
a r t i c l e i n f o a b s t r a c t
Article history: Reliable knowledge of solar radiation is an essential requirement for designing and planning solar energy
Received 14 December 2016 systems. Thus, this paper presents a novel hybrid model for predicting hourly global solar radiation using
Received in revised form 13 January 2017 random forests technique and firefly algorithm. Hourly meteorological data are used to develop the pro-
Accepted 4 February 2017
posed model. The firefly algorithm is utilized to optimize the random forests technique by finding the
Available online 17 February 2017
best number of trees and leaves per tree in the forest. According to the results, the best number of trees
and leaves per tree is 493 trees and one leaf per tree in the forest. Three statistical error values, namely,
Keywords:
root mean square error, mean bias error, and mean absolute percentage error are used to evaluate the
Solar radiation
Prediction
proposed model for the internal and external validation. Moreover, the results of the proposed model
Random forests technique are compared with conventional random forests model, conventional artificial neural network and opti-
ANN mized artificial neural network model by firefly algorithm to show the superiority of the proposed hybrid
Optimization model. Results show that the root mean square error, mean absolute percentage error, and mean bias
Firefly algorithm error values of the proposed model are 18.98%, 6.38% and 2.86%, respectively. Moreover, the proposed
random forests model shows better performance as compared to the aforementioned models in terms
of prediction accuracy and prediction speed.
Ó 2017 Elsevier Ltd. All rights reserved.
1. Introduction models [4] and stochastic algorithm model (Markov chain) [5]
have been developed for modeling global solar radiation.
The availability of solar radiation data in a small scale becomes In general, empirical models are widely used to model global
extremely important for modeling, sizing and controlling of solar solar radiation by correlating the solar radiation with different
energy systems and other applications such as agricultural produc- measured meteorological data and geographical coordinates. How-
tions, hydrological and ecological studies. Solar radiation is usually ever, the accuracy of such models is questionable, especially when
measured using many types of instruments at a particular location. dealing with a highly uncertain data. In the meanwhile, the appli-
These instruments are considered as the most accurate to meaure cation of satellite-derived models is promising for predicting global
the solar radiation data. However, the cost of these instruments solar radiation into a large region using satellite images. This
including calibration and maintenance cost is slightly high. There- method is costly and suffers of the lack of historical meteorological
fore, the availability of these data for all sites is limited in many datasets because it is relatively new. On the other hand, regression
meteorological stations around the world. Moreover, most of the models are good in predicting monthly averages of solar radiation.
recorded datasets of solar radiation suffer of missing data during However, recent challenges to solar energy systems show a dire
outages. Thus, there is a need for prediction more so as to restore need for modeling solar energy systems on hourly basis and conse-
these datasets and to generate solar radiation data series at loca- quently prediction of hourly solar radiation becomes a must for
tion where measurement devices are missing. development in recent research work in the field of solar energy
Based on that, many research works have been done before to [6]. These models have a low performance when modeling global
model the global solar radiation. Several models including the solar radiation on a long term basis, and they are not recom-
empirical models [1,2], satellite-derived model [3], regression mended when there are missing data.
Due the necessity of accurate and reliable modeling of hourly
⇑ Corresponding author.
global solar radiation, machine learning and artificial intelligence
E-mail addresses: [Link]@[Link], [Link]@[Link]
(AI) techniques have been broadly utilized due to their powerful
(I.A. Ibrahim), [Link]@[Link] (T. Khatib). techniques in predicting hourly solar radiation considering the
[Link]
0196-8904/Ó 2017 Elsevier Ltd. All rights reserved.
414 I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425
Nomenclature
Abbreviations
Lb the true label in the training stage of the RFs
RFs random forests
I importance function that got based on the values of Lb
ANN artificial neural network
a the number of samples per leave in tree in the RFs
FFA firefly algorithm
b the number of samples per tree in the RFs
OOB Out-Of-Bag
r ij the distance between any two fireflies i and j at posi-
VI variable importance
tions xi and xj
RMSE root mean squared error
c affixed light absorption coefficient
MAPE mean absolute percentage error
xik , xkj the kth components of the ith and jth firefly positions
MBE mean bias error
within the search-space
d the dimensionality of the problem
Symbols a is a step size scaling factor
bcðtrÞ OBB samples for a specific tree in the RFs b the variation of attractiveness
tr tree number (1, 2, . . . , T) in the RFs b0 the attractiveness at r ¼ 0
T
ðtrÞ ðtrÞ
the total number of trees in the RFs i a randomization parameter
ca , ca;pz the predicted classes for each sample in the tree be- IP m the predicted value
fore and after altering the variable in the RFs Im the target value
xa the sample value in the training stage of the RFs n the number of observations
uncertainty of such data. Artificial neural networks (ANNs) have contain daily global solar radiation, sunshine hours, average, max-
been utilized many times in the literature for this purpose imum and minimum ambient temperature, relative humidity and
[6–10]. A detailed review on solar radiation prediction using ANN water vapor pressure are utilized. The capability of the hybrid
models can be found in [11]. In [12], six ANN-based models to pre- (SVM–WT) is verified by compering its performance with data that
dict daily global solar radiation at a location in Saudi Arabia are are obtained by ANN, genetic programming (GP) autoregressive
presented. A different combination of input variables considering and moving average (ARMA) models.
sunshine duration, ambient temperature, relative humidity and Though, ANNs and SVM are powerful prediction models for
the number of day in the year is utilized for these models. The such a task, there are some novel techniques such as random for-
results show that the sunshine duration and ambient temperature ests technique that can predict global solar radiation in a more
play very important role to gain high accurate results. In [13], a accurate way. Random Forests (RFs) technique is a good option
hybrid model to predict monthly mean daily global solar radiation for predicting hourly global solar radiation. RFs algorithm is a novel
in Saudi Arabia is developed as well. Here, the number of hidden and ensemble machine learning technique which incorporates the
neurons in the ANN model is optimized using particle swarm opti- feature selection. Moreover, RFs have several merits over other
mization (PSO) technique in order to improve model’s accuracy. modeling approaches. In which, RFs can handle with both continu-
The model uses sunshine duration, the number of month in the ous and discrete variables, RFs algorithm does not over fit as a pre-
year, latitude and longitude as inputs of the developed hybrid dictor, and run fast and efficiently when handling large datasets
model. From the results presented in this research, the developed [20]. In addition, RFs algorithm has only two hyper-parameters
hybrid (ANN-PSO) model showed better performance than the which are; the number of trees in the forest, and the number of
conventional ANN model. In [14], an adaptive neuro-fuzzy infer- splits in the subset at each node (number of leaves per tree) [21].
ence system (ANFIS) model is used to predict daily global solar In the meanwhile, some studies are presented for predicting hourly
radiation based on the day of the year at a specific location in and daily solar radiation based on RFs. These studies do not men-
South-Khorasan. The ANFIS is considered as a hybrid model which tion any use of the optimization techniques to optimize the param-
merges the ANNs with the knowledge representation of the fuzzy eters of the RFs (number of trees in the forest and number of leaves
logic. The number of day is the only input used in this model. In per tree). In [22], the authors used daily meteorological data, solar
general, sunshine ratio, ambient temperature as well as relative radiation and air pollution index for three sites in China to develop
humidity are the most correlated meteorological variables to solar a RFs model for predicting daily solar radiation. Here, the default
radiation. Moreover, ANNs are the most powerful tools for predict- number of trees and leaves per tree (500 trees and 5 leaves per
ing hourly solar radiation. These results were supported in [15,16] tree) are used. The results show that the performance of the RFs
as well. model is better than the empirical methodologies (linear, exponen-
Nevertheless, other machine learning techniques are also tial and logarithmic models) in predicting the daily solar radiation.
implemented for this purpose. Recently, support vector machine In addition, in [23], the authors have used a RFs model for predict-
(SVM) technique has been widely used to predict solar radiation ing hourly and daily solar radiation. In this research, the values of
[17]. In [18], a SVM model is employed to predict daily and the number of trees and leaves per tree are not mentioned. More-
monthly global solar radiation in Isfahan, Iran. The radial basis over, the authors in [24] employed different machine learning
function (SVR-rbf), and polynomial basis function (SVR-poly) are techniques to predict hourly solar radiation. In this study, the
utilized for this purpose. Two inputs are used in this research. authors assumed different values of the number of trees and leaves
These inputs are; sunshine hours as well as maximum possible per tree (10, 50, 100 and 300 trees, and 5 and 20 leaves per tree) in
sunshine hours. The superiority of the proposed SVM model has the proposed RFs model. Then, the best values of the number of
been proved as compared to the empirical and the PSO-based mod- trees and leaves are selected based on the performance of the
els. In [19], a new hybrid model using SVM with Wavelet Trans- model.
form (WT) algorithm is presented to predict daily and monthly In general, optimal number of trees and number of leaves per
global solar radiation in Iranian coastal city. Measured data that tree should be found so as to assure accurate prediction of solar
I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425 415
radiation data. Therefore, AI based optimization algorithm can be get a running unbiased estimate of the prediction error as the
employed here as well to support the RFs model. Various optimiza- RFs are constructed in the training phase which is also used to
tion techniques have been implemented to solve the optimization determine the variable importance. The RFs algorithm is a combi-
problems. Classical techniques are used for this purpose which nation of training and testing stages. Fig. 1 shows the main struc-
used a deterministic approach to reach the optimum solution ture of the RFs algorithm.
[25]. However, using the classical techniques for finding the opti- The RFs methodology for regression can be found in [31] which
mum solution is complicated especially when it deals with huge can be described as follows:
search space [26]. In addition, the most of the optimization prob-
lems are highly nonlinear and multimodal, under different com- i. The bootstrap samples are randomly formed from the
plex constraints which became complicated to be solved using learning dataset with replacement. The number of the cre-
these techniques. Currently, artificial intelligence (AI) optimization ated samples equal to the number of the trees. Breiman in
techniques have been extensively employed to solve these com- [31] suggests trying the default (500 trees), half of the
plex optimization problems in numerous domains due to their ease default, and twice the default of the number of trees and
of use, applicability, and global perspective. then taking the best one which there is no any mathemat-
Several AI optimization techniques are used for the purpose of ical formula to set the optimum number of trees into the
optimization such as particle swarm optimization (PSO) [13], arti- forest [33]. In each sample, different patterns of the dataset
ficial bee colony algorithm (ABC) [27], genetic algorithm (GA) [28] are used to develop the RFs model which is grown inde-
and firefly algorithm (FFA) [17]. Most of AI techniques suffer from pendently. Around one-third of these data are left out
partial optimization and fall in local extreme solutions. GA suffers which is called ‘‘out-of-bag” (OOB) data, while the remain-
from low convergence speed, while FFA is the best in terms of con- ing data called in-bag data.
vergence speed. Moreover, FFA has success rate and efficiency that ii. At each node in a regression tree, a number of splits is
make it better than other AI optimization techniques such as par- selected randomly to create the binary rule. However, in
ticle swarm optimization (PSO) and the genetic algorithm (GA) for each split the mean squared error (MSE) is estimated and
continuous and discrete optimization problems [29]. Thus, the fire- compared with the obtained OBB data in order to select
fly algorithm (FFA), which is a novel metaheuristic algorithm, is the best number of splits in each tree which called number
usually used to solve continuous multimodal optimization prob- of leaves per tree. Number of leaves per tree is held constant
lems based on the social behavior of fireflies [30]. during the forest growing which is 5 leaves per tree as the
This study presents a novel hybrid approach by random forests developers of the algorithm set [34].
(RFs) and Firefly Algorithm (FFA) to predict the hourly global solar iii. During the training stage, RFs provide a significant measure
radiation. The Firefly Algorithm (FFA) is applied to determine opti- for variable importance (the individual effects of the inputs
mal RFs parameters. The main objective of the study is to evaluate on the output) on the basis of OOB data and on the
the viability of utilizing the proposed hybrid (RFs–FFA) model for alteration-importance measure. The importance of any vari-
prediction of hourly global solar radiation. To achieve this, hourly able can be obtained by randomly altering all the values of
data measured at a site in Malaysia called ‘‘Klang Valley” are used. this variable (z) in the OOB samples for each tree. The vari-
These hourly data contain sunshine ratio, humidity, ambient tem- able importance measure is calculated as the summation
perature, month number, day number, number of hours per day of the average difference between the prediction accuracy
and global solar radiation. To validate the capability of the pro- before and after altering the variable z over all the predictors
posed RFs–FFA model, the performance of the developed model [31]. Variable importance (VI) decreases when the prediction
is compared with the performance of conventional RFs model accuracy increases. The importance score of each variable is
and conventional ANN model as well as a ANN-FFA model. obtained by the mean importance of all trees by using the
following equation [21]:
P P
2. Proposed hybrid random forests technique based on firefly P xa 2bcðtrÞ
ð ðtrÞ
I Lb ¼ca Þ xa 2bcðtrÞ
ð ðtrÞ
I Lb ¼ca;pz Þ
T cðtrÞ
jb j
jb cðtrÞ
j
algorithm
VIðtrÞ ðzÞ ¼ ð1Þ
T
2.1. Random forests technique
where bcðtrÞ corresponds to OBB samples for a specific tree, tr
represents the tree number (1, 2, . . . , T), T is the total number
Random Forests (RFs) are a bagging based method that used an ðtrÞ ðtrÞ
ensemble machine learning technique of regression trees which of trees, ca and ca;pz are the predicted classes for each sam-
are developed by Breiman (in 2001) [31]. The RFs are a combina- ple in the tree before and after altering the variable, respec-
tion of many decision trees that creates using bootstrapping tech- tively. xa represents the sample value, Lb is the true label;
nique coming from the learning dataset samples of the predictors both are in the training stage, I represents the importance
and choosing randomly at each node. The principle of RFs is done function that got based on the values of Lb , a is the number
with respect to classification and regression trees (CART) model of samples per leave in tree and b is the number of samples
strategy [32]. In the meanwhile, two differences can be noted. First, per tree in the forest.
at each split node in RFs, a learning dataset are randomly selected iv. Then, the outliers in the training dataset are detected by
and the best split is calculated only within this step. Second, no using cluster analysis. In this study, the density model which
clipping step is carried out so all the supposed trees in the forest is one of the cluster analysis types is used. The density model
are maximal trees. can easily detect clusters of points and noise points that do
The RFs rank the input variables by a variable importance mea- not belong to any of these clusters. In machine learning tech-
sure [31], which reflects the impact of the inputs variables on the niques, removing the outlier’s data from the data set will sig-
output based on the prediction accuracy. The RFs algorithm esti- nificantly increase the accuracy of the results. Here, the
mates the importance of the variables comparing the prediction outliers in the training data set are detected, then removed,
error with Out-Of-Bag (OOB) data term. This OOB data term is to and replaced.
416 I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425
v. Finally, a testing stage is done to predict the values. In 2.2. Firefly algorithm (FFA)
the meanwhile, the average of the predictors of all
regression trees are calculated to find the final predicted Firefly algorithm (FFA) is developed by Yang (in late 2007 and
values. 2008) for solving optimization problems [29,35]. The idea of FFA
I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425 417
inspired by the flashing lights that produced by fireflies which are 2.3. Overall structure of the proposed model
short and rhythmic flashing lights. These lights are enabled the
fireflies to attract and assist them in order to find mates, pointe The prediction of solar radiation starts firstly by setting the
to their prey, and protect themselves from their enemies. However, input samples and variables into the RFs algorithm which called
the less bright fireflies are attracted by the bright fireflies [36]. The ‘‘Bagger algorithm” in MATLAB. In this work, the inputs are sun-
flashing lights procedure used by fireflies can be formulated as an shine ratio, humidity, ambient temperature, month number, day
optimization algorithm, since the bright of the flashing lights can number, and number of hours per day. It is noted that there is no
be considered as a coordination with the fitness function to be mathematical formula to set the optimum number of trees [33].
optimized. In the meanwhile, the FFA works based on three rules In this application, the initial number of trees is supposed to be
that can be summarized as: 250 trees as the half of the number of the default tress [31], and
the initial number of leaves per each tree is assumed to be five
Fireflies must be unisex. as it is set by default by the algorithm [34]. These initial values
The less bright firefly is attracted towards the randomly moving of the number of trees and leaves per tree are used just to start
of the brighter fireflies. the parameter tuning which are variable importance measure,
The brightness of each firefly represents the quality of the cluster analysis, and outlier measure.
solution. In general, if the number of trees and leaves per tree is less than
the optimum number, under fitting may occur, and this will lead to
The attractiveness of a firefly is proportionally affected by the
noticed light intensity by the adjacent fireflies. Thus the variation
of attractiveness (b) can be defined by [35],
b ¼ b0 ecr
2
ð2Þ
where xik and xkj are the kth components of the ith and jth firefly
positions within the search-space and d presents the dimensionality
of the problem. The fireflies’ movements are updated during the
searching for the best solution. The movements update is done
based on the current position of the fireflies, the attractiveness
between the fireflies, and a randomization term [29]. The move-
ment of a firefly i is attracted to another brighter firefly j which
updates its position as the following relation [35]:
cr 2ij
xi ðt þ 1Þ ¼ xi ðtÞ þ b0 e ðxj ðtÞ xi ðtÞÞ þ a:i ð4Þ
A representation of a firefly,
An initialization (line 1 in Fig. 2),
A moving operator (line 8 in Fig. 2),
An objective function (lines 2–11 in Fig. 2). Fig. 3. Flowchart of the RFs based on FFA prediction model.
418 I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425
high training and error rate, while over fitting and high variance Fig. 1 are used based on the initial numbers of trees and
may occur when the number of trees and leaves per tree are more leaves per tree. However, the variables and samples tun-
than the optimum number. Thus, a firefly algorithm (FFA) is imple- ing procedure in this stage is described as follows:
mented to optimize the numbers of trees and leaves per tree based
i. The prediction process starts by setting the RFs algo-
on the minimum value of root mean squared error (RMSE) for all
rithm inputs which are the samples and variables. In
the possibilities in the design space. This optimization aims to
this study, the inputs are sunshine ratio, humidity,
improve the performance of the developed RFs model which
ambient temperature, month number, day number,
means the best fitting, less error rate and less variance for the pre-
and number of hours per day, while the output is the
diction function. Fig. 3 shows the RFs prediction flowchart for pre-
global solar radiation.
dicting global solar radiation.
ii. Define the number of tress and number of leaves per
The stages of the proposed RFs algorithm for predicting the
tree. Breiman in [31] suggested to use half of the num-
solar radiation can be described as follows:
ber of the default tress which is 250 trees as the maxi-
mum number of trees into the forest and the developer
of the ‘‘Bagger Algorithm” suggested to use 5 leaves per
Stage I. In this stage, the variables and samples are tuned using
tree as a default [34]. These values are used as initial
the importance variable measure for improving the effi-
values. In this study, the initial numbers of trees and
ciency of the algorithm, and cluster analysis for improving
leaves per tree were 250 trees and 5 leaves per tree,
the accuracy of the results. The first four steps (i–iv) in
respectively.
Fig. 4. Flowchart of firefly algorithm for number of trees and leaves per tree optimization in a RFs prediction model.
I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425 419
iii. Train the RFs algorithm by using the ‘‘Bagger Algo- ii. Create an initial set of fireflies.
rithm” which was developed by MATLAB developer iii. Evaluate firefly parameters for initial population. Here,
using specific settings as a predictor. In this step, the the evaluation is conducted for the attractiveness, dis-
variable importance measure is implemented to find tance and movement of the fireflies.
which of these inputs that will affect more on the out- iv. Run the Bagger Algorithm in the training mode to start
put. The variable importance measure is used to the optimization process for the numbers of trees and
improve the algorithm performance and decrease the leaves per tree.
simulation time. v. Update the position of the fireflies, then check if the
iv. After selecting the important variables, cluster analysis value of the global optima is less than the value of the
is implemented to catch the outliers from the training local optima, if yes then got to step vi, else update the
dataset. Meanwhile, removing the outliers’ data from global optima to be equal to the local optima and go to
the sampled dataset will significantly increase the step iii.
accuracy of the results. In this process, the outliers in vi. Save the best number of trees and the best number of
the training dataset are detected, removed, and leaves per tree.
replaced using the density model. The density model vii. Stop the optimization process and transfer the best num-
defines the clusters as connected dense regions in the bers to the Bagger Algorithm.
dataset space. The data points in the x-axis are pro-
cessed by optics, and those in the y-axis are processed Stage III. Finally, train and test the RFs algorithm using the modi-
by the reachability distance. Thus, an algorithm is devel- fied variables and the optimized parameters of the algo-
oped to remove the outliers and replace them by using a rithm in order to use it to predict the solar radiation for
better way to increase the accuracy of the results. the targeted site. However, the steps (i, ii and v) in
v. Modify the RFs algorithm inputs according to the result Fig. 1 is used in this stage.
of variable importance measure and the replacements
into the training dataset according to the cluster analysis. 2.4. Evaluation of the proposed prediction model
Stage II. In this stage, after modifying the variables and samples,
the number of trees and number of leaves per tree are Three statistical error terms are used to evaluate the results of
optimized using FFA to increase the accuracy of the pre- the RFs model, namely, root mean squared error (RMSE), mean
diction model. The first two steps (i and ii) in Fig. 1 are absolute percentage error (MAPE), and mean bias error (MBE).
used at each trial in the design space to optimize the RMSE is an efficiency indicator of the prediction process; a large
numbers of trees and leaves per tree. Meanwhile, the positive RMSE value represents a large deviation scale in the pre-
procedures in this stage are described as in Fig. 4 and diction values from the target values. MBE is an average deviation
the summarized steps as follows: indicator; a negative value means that the prediction is under-
forecasted and vice versa. MAPE represents an accuracy indicator.
i. Read the objective function. In this study, the objective RMSE, MBE, and MAPE are expressed as follows [38]:
function aims to minimize the RMSE value between the rffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
predicted and actual dataset in the training stage. In 1 Xz
RMSE ¼ ðIm Im Þ ð6Þ
which the objective function as follows: z m¼1
F ¼ min RMSE ð5Þ 1 X z Im IPm
MAPE ¼ m¼1
a ð7Þ
z Im
Fig. 5. Hourly sample of the training dataset of the sunshine ratio for 120 h.
420 I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425
1X
z
MBE ¼ I P Im ð8Þ samples. In which, 8760 hourly sample for each variable. The data-
z m¼1 m set is divided as 70% (42,924 hourly samples) for training and 30%
(18,396 hourly samples) for testing and validation. In the mean-
where IPm represents the predicted value, Im is the target value, and while, the training dataset starts from the 1st of January to the
n is the number of observations. 13th of September (1st-42924th sampling points), while the test-
ing dataset starts from the 14th of September to the 31st of
3. Results and discussion December (42924th-61320th sampling points). Figs. 5–7 show
hourly sample of the training dataset of sunshine ratio, humidity,
To show the effectiveness of the proposed model, a case study and ambient temperature, respectively.
has been presented. Hourly meteorological data for one year for After dividing the dataset, the training dataset with the initial
a site called Klang Valley, Malaysia are obtained. These data are number of trees and number of leaves per tree is used to improve
recorded by the Subang Meteorological Station (latitude: 3.12°N the model performance. The initial number of trees is 250 trees,
and longitude: 101.6°E) are used. The recorded hourly data contain while 5 leaves per tree are used as a default of the algorithm as
sunshine ratio, humidity, ambient temperature, month number, the suggestion of the developers. The first step is conducted using
day number, and number of hours per day. These data are used the variable importance measure to improve the accuracy of the
as inputs for the proposed model, while hourly global solar model. In case that some of the used inputs are not contributing
radiation is used as an output. The dataset contains 61,320 hourly to the response, then the accuracy of the model will decrease
Fig. 6. Hourly sample of the training dataset of the humidity for 120 h.
Fig. 7. Hourly sample of the training dataset of the ambient temperature for 120 h.
I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425 421
and the simulation time will increase. As it is mentioned, 6 inputs replaced. Outlier detection is the identification of observations that
are used for predicting the hourly global solar radiation. It is clear do not conform to a prospective pattern in a dataset. It is normally
from Fig. 8 that all of the inputs are contributing to the output but accomplished with statistics and thresholds. Cluster analysis is one
in different amounts. Therefore, Fig. 8 shows the variable impor- of the most popular technique for detecting outliers that not
tance rates for the seven inputs used in this study. From the figure, belong to the dataset. A normal distribution model is used to ana-
the most important variable is the number of the month with a lyze the data set. The outliers expected in the data set are depicted
rate of 3.2 out of 3.5. Second, sunshine ration which has a rate of in Fig. 9, which shows the number of observations that detected to
3 out of 3.5, whereas the hour has a rate of 2.6 out of 3.5. In the be outliers. The first pattern contains 45.6% of the outliers, 44% in
meanwhile, the rates of the ambient temperature and humidity the second pattern, 5.7% in the third pattern, 0% in all of the fourth,
are 1.4 and 1.25 out of 3.5, respectively. Finally, number of the fifth, sixth, seventh, eighth and tenth patterns, while it is 2.4% in
day has a rate of 0.6 out of 3.5. Here, there is no input has a rate the ninth pattern. However, the algorithm is developed to remove
of 0 out of 3.5. This mean that no any input will be removed. these outliers and replace them by using a better way to increase
In machine learning techniques, removing the outliers from the the accuracy of the results.
dataset will significantly increase the accuracy of the results. Here, After measuring the variable importance and replacing the out-
the outliers in the training dataset are detected, then removed and liers in the training dataset using the initial numbers of trees and
Fig. 10. Convergence characteristic curve for FFA in optimizing the number of trees and leaves per tree in the RFs model based on the minimum RMSE.
leaves per tree (250 trees and 5 leaves per tree), an optimization of The convergence curve of the FFA for optimizing the numbers of
these numbers is conducted using FFA. This optimization aims to trees and leaves per tree has been shown in Fig. 10, where the x-
find the best numbers of trees and leaves per tree based on the axis is the number of iterations and the y-axis is the fitness value
minimum value of RMSE. This optimization will significantly of the benchmark function (RMSE). According to the figure, the
improve the performance of the proposed model by increasing FFA maintains a higher convergence rate, and it becomes plunged
the accuracy of the predicted results. Here, an initial population into local optima after a bout 30 iterations. It is obvious that FFA
of the fireflies is set randomly. The population size is assumed to displays a faster convergence rate to find the final global optima.
be 100 and the maximum number of iterations is 500. The FFA Based on that, the best numbers of trees and leaves per tree are
parameters are set as; the step size scaling factor (a), the variation 493 trees and 1 leave per tree.
of attractiveness (b), and the affixed light absorption coefficient The RFs based FFA model is used to predict the hourly global
(c) assumed as mentioned in [39] that to be 0.2, 1 and 1, solar radiation. Using the best numbers of trees and leaves per tree
respectively. that obtained by FFA (493 trees and 1 leaf), 70% of the dataset in
Fig. 11. Hourly sample of the predicted solar radiation by RFs-FFA model.
I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425 423
Fig. 12. Convergence characteristic curve for FFA in optimizing the number of neurons in the ANN model based on the minimum RMSE.
Table 1 the location which mentioned above are used to train the proposed
The internal parameters of the RFs, RFs-FFA, ANN and ANN-FFA prediction models. model. Then, 30% of the dataset are used for testing and validating
Index RFs RFs-FFA the model. Based on that, a sample of the predicted solar radiation
using RFs-FFA model is illustrated in Fig. 11.
No. of trees 500 493
No. of leaves 5 1 To show the superiority of the proposed model, the predicted
ANN ANN-FFA global solar radiation is compared with that solar radiation data
No. of neurons 3 6 which have been predicted by conventional RFs and ANN as well
as optimized ANN based on FFA models. Firstly, RFs model with
Fig. 13. Hourly sample of the predicted solar radiation by ANN, ANN-FFA, RFs and RFs-FFA models.
424 I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425
Table 2
Statistical values and time consumption of solar radiation prediction techniques
Model RMSE (W/m2) RMSE (%) MAPE (%) MBE (W/m2) MBE (%) Consumption time (s)
ANN 95.9144 26.4459 17.1587 25.3475 6.9889 7.8010
ANN-FFA 85.1201 23.4698 13.0014 18.0345 4.9725 7.9915
RFs 74.4561 20.5293 9.7894 13.1095 3.6146 7.2801
RFs-FFA 68.8359 18.9797 6.3826 10.3838 2.8631 6.4885
the default number of trees and leaves (500 trees and 5 leaves per tial solutions for the numbers of trees and leaves per tree are gen-
tree) is used in this validation. In this model, the flowchart that is erated randomly in the RFs and the improvement is carried out by
illustrated in Fig. 3 is implemented without any optimization of the FFA which tries to optimize these numbers in the RFs. The pre-
model’s parameters (the so called conventional RFs mode). Sec- dicted data by RFs-FFA model is compared with that predicted by
ondly, the ANN model which is presented in [16] is also utilized ANN, ANN-FFA and RFs models to show the superiority of the pro-
in this validation (feed-forward ANN topology). According to the posed model. Three statistical error values, namely, RMSE, MAPE,
literature, the number of hidden neurons cannot be obtained with- and MBE are employed to evaluate the accuracy of the proposed
out conducting model’s training and hen estimating the general- model. On the basis of the results, the RFs-FFA model is found to
ization error. This is to say that there is no direct rule to get the be accurate in predicting the hourly global solar radiation than
optimum number of neurons in the hidden layer [40]. Based on the aforementioned models. The RMSE, MAPE and MBE values for
that, the number of hidden neurons in the utilized ANN model the hybrid RFs-FFA model are around 18.98%, 6.38% and 2.86%,
using a ‘‘trial and error” process is found to be 3 neurons. respectively. Based on that, the proposed RFs-FFA model can there-
Here, if the number of the hidden neurons is lower than the fore be adjudged an efficient and appealing machine learning
optimum, an under fitting may occur, and this may lead to a high based model to obtain high level of accuracy in predicting the
generalization error. On the other hand, if the number of the hid- hourly global solar radiation. Thus, the RFs-FFA model is recom-
den neurons is higher than the optimum, an over fitting and high mended for predicting the global solar radiation.
variance may occur [40]. Thus, the FFA is used to optimize the
number of the hidden neurons. Here, the optimization function
aims to minimize the RMSE. Accordingly, an initial population of
References
the fireflies has been set randomly. Meanwhile, the size of the pop-
ulation is assumed to be 100 and the maximum number of itera- [1] Halawa E, GhaffarianHoseini A, Li D Hin Wa. Empirical correlations as a means
tions is 500. The FFA parameters are set as; a equal to 0.2, b for estimating monthly average daily global radiation: a critical overview.
equal to 1 and c equal to 1 [39]. Renew Energy 2014;72:149–53. [Link]
renene.2014.07.004.
The convergence curve of the FFA for optimizing the numbers of [2] Besharat F, Dehghan AA, Faghih AR. Empirical models for estimating global
the hidden neurons is shown in Fig. 12. According to the figure, the solar radiation: a review and case study. Renew Sustain Energy Rev
local optima has been obtained after about 20 iterations. As a 2013;21:798–821. [Link]
[3] Viana TS, Rüther R, Martins FR, Pereira EB. Assessing the potential of
result, the best number of the hidden neurons is 6. The internal concentrating solar photovoltaic generation in Brazil with satellite-derived
parameters of the RFs-FFA, RFs, ANN and ANN-FFA models are direct normal irradiation. Sol Energy 2011;85:486–95. [Link]
summarized in Table 1. 10.1016/[Link].2010.12.015.
[4] Kumar R, Aggarwal RK, Sharma JD. Comparison of regression and artificial
Using the same case study data and the percentages of the
neural network models for estimation of global solar radiations. Renew Sustain
training and testing dataset, the RFs, ANN and ANN-FFA models Energy Rev 2015;52:1294–9. [Link]
are used to predict the hourly global solar radiation as benchmarks [5] Hocaoĝlu FO. Stochastic approach for daily solar radiation modeling. Sol
Energy 2011;85:278–87. [Link]
for the external validation of the proposed model. A sample of the
[6] Khatib T, Mohamed A, Sopian K. A review of solar energy modeling techniques.
predicted results by ANN, ANN based on FFA, RFs and RFs based Renew Sustain Energy Rev 2012;16:2864–9. [Link]
FFA models are shown in Fig. 13. rser.2012.01.064.
Based on Fig. 13, the results clearly indicate that the hybrid [7] Mellit A, Kalogirou SA. Artificial intelligence techniques for photovoltaic
applications: a review. Prog Energy Combust Sci 2008;34:574–632. [Link]
model (RFs-FAA) has performed better than the ANN, ANN-FFA [Link]/10.1016/[Link].2008.01.00.
and RFs models. Moreover, the predicted values by RFs-FFA model [8] Benson RB, Paris MV, Sherry JE, Justus CG. Estimation of daily and monthly
are closer to the actual data as compared to results obtained by the direct, diffuse and global solar radiation from sunshine duration
measurements. Sol Energy 1984;32.
other models. Using the evaluation criteria that mentioned in sec- [9] Bakirci K. Models of solar radiation with hours of bright sunshine: a review.
tion 2.4, the statistical values are calculated as shown in Table 2. Renew Sustain Energy Rev 2009;13:2580–8. [Link]
From Table 2, the RMSE, MAPE and MBE are about 18.98%, 6.38% rser.2009.07.011.
[10] Jebaraj S, Iniyan S. A review of energy models. Renew Sustain Energy Rev
and 2.86%, respectively. These results show the superiority of the 2006;10:281–311. [Link]
proposed method as compared to the other considered methods. [11] Yadav AK, Chandel SS. Solar radiation prediction using Artificial Neural
Furthermore, the RFs-FFA model is faster than the ANN, ANN-FFA Network techniques: a review. Renew Sustain Energy Rev 2014;33:772–81.
[Link]
and RFs models in terms of training and testing processes by
[12] Benghanem M, Mellit A, Alamri SN. ANN-based modelling and estimation of
1.3125 s., 1.503 s. and 0.7916 s., respectively. These results prove daily global solar radiation data: a case study. Energy Convers Manage
the superiority of RFs-FFA model in predicting the hourly global 2009;50:1644–55. [Link]
[13] Mohandes MA. Modeling global solar radiation using Particle Swarm
solar radiation compared with the ANN, ANN-FFA and RFs models.
Optimization (PSO). Sol Energy 2012;86:3137–45. [Link]
[Link].2012.08.005.
4. Conclusion [14] Mohammadi K, Shamshirband S, Tong CW, Alam KA, Petković D. Potential of
adaptive neuro-fuzzy system for prediction of daily global solar radiation by
day of the year. Energy Convers Manage 2015;93:406–13. [Link]
In this study, a novel hybrid (RFs-FFA) model is developed for 10.1016/[Link].2015.01.02.
predicting the hourly global solar radiation for a location in Malay- [15] Khatib T, Elmenreich W. Modeling of photovoltiac systems using Matlab:
sia. Hourly meteorological data for one-year period are used for Simplified green codes. Wiley; 2016.
[16] Khatib T, Mohamed A, Sopian K, Mahmoud M. Assessment of artificial neural
developing the hybrid model. The FFA is implemented to optimize networks for hourly solar radiation prediction. Int J Photoenergy 2012 2012.
the numbers of trees and leaves per tree in the RFs technique. Ini- [Link]
I.A. Ibrahim, T. Khatib / Energy Conversion and Management 138 (2017) 413–425 425
[17] Olatomiwa L, Mekhilef S, Shamshirband S, Mohammadi K, Petković D, Sudheer [27] Luo J, Liu Q, Yang Y, Li X, Cao MCW. An artificial bee colony algorithm for
C. A support vector machine-firefly algorithm-based model for global solar multi-objective optimisation. Appl Soft Comput 2017;50:235–51. [Link]
radiation prediction. Sol Energy 2015;115:632–44. [Link] [Link]/10.1016/[Link].2016.11.014.
[Link].2015.03.015. [28] Aybar-Ruiz A, Jiménez-Fernández S, Cornejo-Bueno L, Casanova-Mateo C,
[18] Mohammadi K, Shamshirband S, Anisi MH, Amjad Alam K, Petković D. Support Sanz-Justo J, Salvador-González P, et al. A novel grouping genetic algorithm-
vector regression based prediction of global solar radiation on a horizontal extreme learning machine approach for global solar radiation prediction from
surface. Energy Convers Manage 2015;91:433–41. [Link] numerical weather models inputs. Sol Energy 2016;132:129–42. [Link]
enconman.2014.12.015. org/10.1016/[Link].2016.03.015.
[19] Mohammadi K, Shamshirband S, Tong CW, Arif M, Petkovi D, Sudheer C. A new [29] Yang XS. Firefly algorithms for multimodal optimization. In: Stochastic
hybrid support vector machine-wavelet transform approach for estimation of algorithms: foundations and applications. Springer; 2009.
horizontal global solar radiation. Energy Convers Manage 2015;92:162–71. [30] Yang XS. Firefly algorithm in engineering optimization. John Wiley & Sons Inc.;
[Link] 2010.
[20] Tooke TR, Coops NC, Webster J. Predicting building ages from LiDAR data with [31] Breiman L. Random forests. Mach Learn 2001;45:5–32. [Link]
random forests for building energy modeling. Energy Build 2014;68:603–10. 10.1023/A:1010933404324.
[Link] [32] Breiman L, Friedman J, Olshen R, Stone C. Classification and regression
[21] Guo L, Chehata N, Mallet C, Boukir S. Relevance of airborne lidar and trees. Monterey, CA: Wadsworth and Brooks; 1984.
multispectral image data for urban scene classification using Random [33] Liaw A, Wiener M. Classification and regression by random forest. R News
Forests. ISPRS J Photogramm Remote Sens 2011;66:56–66. [Link] 2002;2:18–22. [Link]
10.1016/[Link].2010.08.007. [34] MATLAB. TreeBagger. Mathworks; 2016. <[Link]
[22] Sun H, Gui D, Yan B, Liu Y, Liao W, Zhu Y, et al. Assessing the potential of stats/[Link]> [accessed August 28, 2016].
random forest method for estimating solar radiation using air pollution index. [35] Yang XS. Nature-inspired metaheuristic algorithms. 2nd ed. UK: Luniver Press;
Energy Convers Manage J 2016;119:121–9. [Link] 2008.
enconman.2016.04.051. [36] Apostolopoulos T, Vlachos A. Application of the firefly algorithm for solving the
[23] Kratzenberg M, Helmut H, Preede P, Georg H. Identification and handling of economic emissions load dispatch problem. Int J Comb 2011. [Link]
critical irradiance forecast errors using a random forest scheme – a case study 10.1155/2011/52380.
for southern Brazil. Energy Proc 2015;76:207–15. [Link] [37] Yang XS, Xingshi H. Firefly algorithm: recent advances and applications. Int J
[Link].2015.07.900. Swarm Intell 2013:36–50.
[24] Gala Y, Fernández Á, Díaz J, Dorronsoro JR. Hybrid machine learning [38] Ameen AM, Pasupuleti J, Khatib T, Elmenreich W, Kazem HA. Modeling and
forecasting of solar radiation values. Neurocomputing 2015:1–12. [Link] characterization of a photovoltaic array based on actual performance using
[Link]/10.1016/[Link].2015.02.078. cascade-forward back propagation artificial neural network. J Sol Energy Eng
[25] Deb K, Beyer H. Real coded genetic algorithms with simulated binary 2015;137:1–5. [Link]
crossover: studies on multi modal and multi objective problems. Complex [39] Shareef H, Ibrahim AA, Mutlag AH. Lightning search algorithm. Appl Soft
Syst 1995:431–54. Comput 2015;36:315–33. [Link]
[26] Kanzow C, Yamashita N, Fukushima M. Levenberg – Marquardt methods with [40] Geman S, Doursat R, Bienenstock E. Neural networks and the biad/variance
strong local convergence properties for solving nonlinear equations with dilemma. Neural Comput 1992;4:1–58. [Link]
convex constraints. J Comput Appl Math 2004;172:375–97. [Link] neco.1992.4.1.1.
10.1016/[Link].2004.02.01.