Received August 26, 2021, accepted September 13, 2021, date of publication September 22, 2021, date of current
version October 1, 2021.
Digital Object Identifier 10.1109/ACCESS.2021.3114809
Improving Stock Price Prediction Using
Combining Forecasts Methods
MOHAMMAD RAQUIBUL HOSSAIN 1,2 , MOHD TAHIR ISMAIL 1,
AND SAMSUL ARIFFIN BIN ABDUL KARIM 3
1 Schoolof Mathematical Sciences, Universiti Sains Malaysia, Pulau Pinang 11800, Malaysia
2 Department of Applied Mathematics, Noakhali Science and Technology University, Noakhali 3814, Bangladesh
3 Fundamental and Applied Sciences Department, Centre for Systems Engineering (CSE), Institute of Autonomous System, Universiti Teknologi PETRONAS,
Bandar Seri Iskandar, Seri Iskandar, Perak 32610, Malaysia
Corresponding author: Samsul Ariffin Bin Abdul Karim (samsul_ariffin@[Link])
This work was supported by Universiti Teknologi PETRONAS (UTP) and the Petroleum Research Fund (PRF) through Research Grant
YUTP-FRG: 015LC0-315.
ABSTRACT This study presents an outcome of pursuing better and effective forecasting methods. The study
primarily focuses on the effective use of divide-and-conquer strategy with Empirical Mode Decomposition
or briefly EMD algorithm. We used two different statistical methods to forecast the high-frequency EMD
components and the low-frequency EMD components. With two statistical forecasting methods, ARIMA
(Autoregressive Integrated Moving Average) and EWMA (Exponentially Weighted Moving Average),
we investigated two possible and potential hybrid methods: EMD-ARIMA-EWMA, EMD-EWMA-ARIMA
based on high and low-frequency components. We experimented with these methods and compared their
empirical results with four other forecasting methods using five stock market daily closing prices from
the S&P/TSX 60 Index of Toronto Stock Exchange. This study found better forecasting accuracy from
EMD-ARIMA-EWMA than ARIMA, EWMA base methods and EMD-ARIMA as well as EMD-EWMA
hybrid methods. Therefore, we believe frequency-based effective method selection in EMD-based hybridiza-
tion deserves more research investigation for better forecasting accuracy.
INDEX TERMS EMD, combining forecasts, time-series analytics, ARIMA, EWMA.
I. INTRODUCTION Average (ARIMA) is still widely used for its satisfactory
In many actual events, individuals or organizations require effectiveness. Some recent ARIMA based works encom-
making decisions that involve some certainty and uncer- pass [1]–[4]. Smoothing methods are also efficient and much
tainty. That is why we investigate deterministic, probabilistic, used till today, and some recent research studies include
or mixed scientific methods for a better decision. Some events [5]–[7]. These comparatively basic and traditionally useful
are related to time series (data series indexed with time order). methods serve up to general-purpose expectation level of
Avid practitioners search for effective and robust methods forecasting accuracy. However, many researchers are contin-
to make a better decision for important time series based ually investigating to find better and more effective methods
on future events. Nevertheless, projecting future reality or due to the limitations of benchmark methods and expecta-
forecasting future events is quite challenging due to method- tion for more accuracy. Some of them involve fuzzy theory-
ological limitations and unexpected uncertainty. Therefore, based modeling [8], [9]; some are hybridization with neural
the research study’s scope of the present paper is to find networks and machine learning methods [10]–[15]. With the
improved or better forecasting methods to outperform exist- current popularity in non-time-series applications, machine
ing or benchmark methods. learning or artificial neural networks methods got adequate
Indeed, there are effective benchmark methods with their practical attention and research interest.
different variants. The Autoregressive Integrated Moving However, these methods fail to win the race of fore-
casting competitions [16], [17]; they also have some draw-
The associate editor coordinating the review of this manuscript and backs, including sophistication and high computational cost.
approving it for publication was Frederico Guimarães . In quest of better forecasting methods, some research studies
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
VOLUME 9, 2021 132319
M. R. Hossain et al.: Improving Stock Price Prediction Using Combining Forecasts Methods
focus on hybrid forecasting methods based on Empirical following form:
Mode Decomposition (EMD), a locally adaptive data decom-
yt = l + θ1 yt−1 + θ2 yt−2 + . . . + θp yt−p + st
position algorithm which was originally devised for signal
processing. Thus, this research study focuses on EMD based +ϕ1 st−1 + ϕ2 st−2 + . . . + ϕq st−q (1)
hybrid forecasting with two prominently useful statistical where l is constant; yi ’s are past p autoregressive terms; sj ’s
forecasting methods ARIMA and Exponentially Weighted are past q error terms or random shocks; θi ’s and ϕj ’s are
Moving Average (EWMA). The EMD-ARIMA-EWMA or parameters. If the order of differencing (also equivalent to
EMD-EWMA-ARIMA or both methods have high potential- the order of integration) d is required for nonstationary series
ity to be very useful because of the theoretical combination X to make it stationary Y , ARIMA is better written with
from EMD components. notation ARIMA (p, d, q).
The concept of EMD is deeply rooted from signal pro-
cessing and, therefore, widely used in that field. Two cru- B. EWMA
cial original contributions of EMD in time series were EWMA (also known as single or simple exponential smooth-
the works of [18] (the basic and foundational resource) ing) is presented in [6] and [7]. EWMA produces smoothing
and [19] (application in the financial domain). Besides sig- series S = {S2 , S3 , . . . , St , . . .} from an original series X =
nal processing, EMD based research has broad application {x1 , x2 , x3 , . . . , xt , . . .}. Started with S2 = x1 as a seed value,
areas for analyzing and forecasting time series data in var- recursive relation of EWMA has the following form:
ious fields. Some recent EMD based forecasting studies
include [11]–[15], [20], [21]. Among others, these stud- S2 = x1 (2)
ies reflect the effectiveness of EMD-based hybridization St = ωxt−1 + (1 − ω) St−1 , 0 < ω < 1, t ≥ 3, (3)
methods.
where ω is smoothing parameter, t is time order and St is
EMD is efficient in adaptive decomposition, while sta-
present smoothing term found from the convex combination
tistical forecasting methods ARIMA and EWMA also are
of most recent original data xt−1 and most recent smoothing
simple but effective. We hypothesized that all these three
data St−1 . EWMA’s abbreviated form is:
approaches’ effective synergy could produce better fore-
t−2
casting results for nonlinear nonstationary mixed-frequency X
time series. Using statistical forecasting methods ARIMA St = ω (1 − ω)i−1 xt−i + (1 − ω)t−2 S2 , t ≥ 3 (4)
and EWMA with EMD based on high-frequency and low- i=1
frequency EMD components, this research study investigates The weights ω (1 − ω)t for past data geometrically or
forecasting strategy improvement, which is the aim of the exponentially diminishes as (1 − ω)t decrease with the
study. We employed five stock data sets of Toronto Stock increase of t.
Exchange based S&P/TSX 60 Index to experiment with the
effectiveness of proposed EMD-ARIMA-EWMA and EMD- C. EMD-BASED HYBRIDIZATIONS
EWMA-ARIMA methods. Also, their forecasting perfor- [18] proposed that EMD is an integral part of Hilbert-Huang
mances were compared with benchmark and hybrid methods transforms for signal processing analysis purposes. From the
using both absolute and relative error measures. practical viewpoint, EMD is an adaptive decomposition algo-
The following section presents statistical benchmark rithm which upholds local features. Analysis or decomposi-
methods and their EMD-based hybridizations. Other sections tion of EMD signal data reserves time domain. The principal
subsequently convey proposed experimental methods, exper- procedure of EMD, known as the sifting process, uses local
imental results, discussion for study outcome and conclusion. modal values of a time series to produce an orthogonal set of
sub-signals or equivalently sub-series of different amplitude
II. THE TWO STATISTICAL BENCHMARK METHODS AND and frequency. These orthogonal subseries commonly called
THEIR EMD-BASED HYBRIDIZATIONS Intrinsic Mode Functions (IMFs) must qualify some criteria.
This section mainly and concisely describes two selected The EMD sifting process is repeated until all potential IMFs
statistical benchmark methods ARIMA and EWMA, extraction is completed. The final remnants of the whole
and their EMD-based hybridizations EMD-ARIMA and process are called the residual. An IMF is defined with two
EMD-EWMA. qualifying criteria. One criterion should have a zero mean
value, i.e., it is an oscillatory function that oscillates around
A. ARIMA zero. Another criterion is the difference between the number
The concept of Autoregressive Moving Average (ARMA) [2], of extrema, and the number of zero-crossings will be zero
[4], which is applicable for stationary time series, comes or one but not beyond that. In the sifting process, EMD first
before ARIMA concept of nonstationary series. Generally, produces comparatively high frequency or rapidly oscillatory
a nonstationary series X = {x1 , x2 , x3 , . . . , xt , . . .} can be IMFs and gradually extracts comparatively low frequency or
transformed into stationary Y = {y1 , y2 , y3 , . . . , yt , . . .} slowly oscillating IMFs and residue finally. If N is the total
by taking lag difference once or twice, known as the quantity of data, the maximum quantity of IMFs n is less than
order of differencing. An ARMA (p, q) method has the log2 (N ).
132320 VOLUME 9, 2021
M. R. Hossain et al.: Improving Stock Price Prediction Using Combining Forecasts Methods
EMD sifting process applied in IMFs extraction from a EMD based hybridizations involve a significant part in time
series or signal data set x (t) includes the following steps: series-based research domains. Both EMD-ARIMA (focused
Step 1: To find local extrema (i.e., minima and maxima) of in [22] and [23]) and EMD-EWMA hybrid methods have
x (t) or temporary remainder (remainder found immediately similar algorithms. The first step of these methods involves
after each IMF collection). decomposing the given series or signal using EMD sift-
Step 2: To fit and formulate two cubic spline envelopes ing process to extract EMD components. The second step
with local extrema-based two sub-datasets. One envelope involves fitting and forecasting the already obtained EMD
namely upper envelope (uE ) is for data above local mean, component using a method (here ARIMA or EWMA). In the
i.e., almost half of all data. Another envelope known as lower third step, all the forecasted point data or subseries are aggre-
envelope (lE ) fits data below local mean, i.e., remaining half gated to produce final results, and errors are calculated from
of dataset. this final resultant series data and test data.
Step 3: To obtain arithmetic mean of uE and lE as mean
envelope mE . III. DESCRIPTION OF PROPOSED
uE + lE EXPERIMENTAL METHODS
mE = (5) This section describes two experimental methods and error
2
measures used for performance comparison.
Step 4: mE is subtracted from x (t) and 1st temporary
remainder H1 is found.
A. EXPERIMENTAL METHODS
x (t) − mE1 = H1 (6) This study investigated the frequency-based method selec-
tion for EMD components. Firstly, EMD sifting algorithm
Step 5: Check if H1 follows IMF characteristic defini-
was applied to obtain EMD components, i.e., IMFs along
tion. If it does, then it is an IMF and follows step 6. If H1
with residual. These components were classified into two
does not satisfy IMF definition, the process for next mean
categories. One category, say category I, is a high-frequency
envelope and temporary remainder finding is done for it
category that involves rapidly oscillating stationary IMFs;
following steps 2 to 4 which subsequently produces other
another category, category II, is the low-frequency category
mean envelopes (from 2nd to higher) and relevant temporary
involving the IMFs from slowly oscillating nonstationary to
remainders.
residual (refer to Fig. 2). For considering EMD components
H1 − mE2 = H2 , H2 −mE3 = H3 , . . . ,Hk−1 −mEk = Hk in these frequency categories, the Augmented Dickey-Fuller
(7) (ADF) test was used. For any EMD component, if the p-value
of ADF test is less than 5% or 0.05, it fails to accept the
Stopping criterion like Cauchy Convergence Crite- null hypothesis that the EMD component possesses unit root
rion (CCC) involving standard deviation is applied in IMF (or in other words it is non-stationary). Also, it cannot reject
sifting process, which is defined as: the null hypothesis if the p-value is more significant than
T 0.05. Therefore, when the p-value of any EMD component
X (Hk−1 (t) − Hk (t))2 is smaller than 0.05, it is stationary or category I EMD com-
CCCk = 2 (t)
(8)
t=0
Hk−1 ponent. Two statistical methods ARIMA and EWMA were
used to fit and forecast EMD components in these categories,
IMF sifting process is terminated when CCCk becomes
which led to two experimental methods. One method is EMD-
smaller than a pre-set minimum value.
ARIMA-EWMA, where stationary components were fitted
Step 6: If an Hk satisfies stopping criterion, it is considered
and forecasted using ARIMA while remaining non-stationary
and selected as an IMF. Let 1st IMF be F1 = Hk . Now, a con-
components were fitted and forecasted with EWMA. Another
tinual remainder is calculated by subtracting F1 from x (t) .
method EMD-EWMA-ARIMA applies EWMA for station-
Thus, 1st continual remainder is found as R1 = x (t) − F1 .
ary and ARIMA for nonstationary components.
Likewise, following steps 1 to 5 current continual remainder
The experimental method I: EMD-ARIMA-EWMA
Ri , other next continual remainders Ri+1 ’s are calculated
Let EMDcomi ∈ {IMF1 , IMF2 , . . . ., IMFi , . . . , IMFn ,
during subsequent IMFs extraction.
residual}, i ∈ {1, 2, 3, . . . ., n + 1} and i ∈ N .
R2 = R1 − F2 , R3 = R2 − F3 , . . . ,Rn = Rn−1 − Fn (9)
forecast1 (EMDcomi )
Continual remainders are repeatedly obtained until the last
ARIMA (EMDcomi ) ,
one represents a monotonic function or a monic of first-
if p − value (EMDcomi ) ≤ 0.05
degree function. The last or final continual remainder Rn of
= (10)
the process is known as residual.
EWMA (EMDcomi ) ,
Even though EMD or Hilbert-Huang transforms were ini-
if p − value (EMDcomi ) > 0.05
tially developed for signal analysis purposes, many research
EMD − ARIMA − EWMAforecast
studies of other fields, including economics and finance, used X
it widely to analyze and forecast time series. Accordingly, = forecast1 (EMDcompi ) (11)
VOLUME 9, 2021 132321
M. R. Hossain et al.: Improving Stock Price Prediction Using Combining Forecasts Methods
Experimental method II: EMD-EWMA-ARIMA Algorithm EMD-ARIMA-EWMA
Input: X , a pre-processed stock price dataset.
forecast2 (EMDcomi ) Output: Error results from EMD-ARIMA-EWMA fore-
cast.
EWMA (EMDcomi ) ,
Step 1: Begin
if p − value (EMDcomi ) ≤ 0.05
= (12) Step 2: Read X
ARIMA (EMDcom i) , Step 3: Split X into Xtrain and Xtest and store; set future
if p − value (EMDcomi ) > 0.05 horizon h
EMD − EWMA − ARIMAforecast Step 4: Apply EMD sifting process on Xtrain ; store IMFs
X IMF1 , IMF2 , . . . , IMFi , . . . , IMFn and residual
= forecast2 (EMDcomi ) (13) Step 5: If IMFi is stationary or category I component,
fit and forecast with ARIMA method, i.e., ARIMA (IMFi )
Theoretically, it is expected that suitable model selec- and forecast (ARIMA (IMFi ) , h); if IMFi or residual is
tion and parameter value computation of any method (here nonstationary, fit and forecast with EWMA method,
ARIMA and EWMA) can fit data optimally to produce better i.e., EWMA (IMFi ) and forecast (EWMA (IMFi ) , h)
forecast data. Hence, if two different suitable and effec- Step 6: Aggregate all forecast data, i.e.,
tive methods are used for different types or categories of
forecast (ARIMA (IMFi ) , h)
i
EMD components, the new combined method is expected to X
aggrforecast = +forecast (EWMA (IMFi ) , h)
improve the forecast results of two individual methods. Since
EMD is a locally adaptive data-driven approach, if EMD 1 +forecast (EWMA (residual) , h)
components are efficiently fitted and forecasted with optimal Step 7: Define error (M , N )
Step 8: Compute error, err = error Xtest , aggrforecast
models, the final aggregate result is expected to be superior
to its constituents. In EMD-ARIMA-EWMA, ARIMA and Step 9: Print error err and output aggrforecast for EMD-
EWMA were optimized by information criterion, e.g., AIC ARIMA-EWMA method
and optimal parameter selection (also parameter values come Step 10: End
from optimization approaches with the underlying methods),
which is in support of principle of parsimony and optimum
model selection to avoid under-fitting or over-fitting. How- Xtest and forecasted data aggrforecast . In step 9, error results
ever, the out-sample forecast errors or accuracy of a method and forecasted data are given as output.
are better indicators for its outperformance than compared
methods. C. ERROR TOOLS FOR COMPARING FORECASTING
PERFORMANCE
B. ALGORITHMS OF EXPERIMENTAL METHODS This research study used two absolute error measuring tools
Algorithms of the two experimental methods EMD-ARIMA- and one relative error tool to compare the forecast perfor-
EWMA and EMD-EWMA-ARIMA are very similar. mance of employed methods. The absolute error tools are
Therefore, below is presented algorithm only for EMD- Root Mean Squared Error (RMSE) and Mean Absolute Error
ARIMA-EWMA method. (MAE). Also, the relative error tool is Mean Absolute Per-
A brief explanation of EMD-ARIMA-EWMA algorithm centage Error (MAPE). These error formulae are defined in
is presented here. In step 2, input data X is read and stored (14-16):
(here 1416 data in each dataset). Then step 3 X is split into v
u n 2
uP
training subset Xtrain (here 1400 data) and test subset Xtest u Dt
(16 data). In step 4, EMD algorithm is applied to the training t 1
subset Xtrain and EMD components are computed and stored RMSE = (14)
n
(here 9 IMFs and residual). In step 5, EMD components Pn
are tested for whether it is stationary or nonstationary using |Dt |
1
ADF test. An EMD component is fitted and forecasted using MAE = (15)
ARIMA method if it is stationary; while an EMD component n
n
P Dt
is nonstationary, it is fitted and forecasted using EWMA. yt
In contrast, it is in the opposite direction for EMD-EWMA- 1
MAPE = , (16)
ARIMA. In step 6, all component forecasts are aggregated n
to obtain complete final forecast data. In step 7, an error where Dt = yt − ft , the difference of forecast data from
function is defined with arbitrary variable M and N which original test data; t is time sequence, and n is the quantity of
will respectively take test data and forecasted data as input test data (or forecast data). If the value of error tools for any
and produce three types of error or accuracy results which method is smaller, then the forecasting method is considered
are defined in next sub-section. In step 8, the errors of the better, i.e., when a method produces immense error value, its
method are computed using the error function with test data forecasting performance is poor.
132322 VOLUME 9, 2021
M. R. Hossain et al.: Improving Stock Price Prediction Using Combining Forecasts Methods
IV. EXPERIMENTAL RESULTS
This section presents experimental data sets (with summary
statistics and graphs) and forecast accuracy results of all
methods used in this study (with brief description on com-
putation time).
A. DATASETS
Data sets used for experimenting the proposed EMD-
ARIMA-EWMA method are closing price data of five stock
listed companies namely Bank of Montreal (BMO), Bank
of Nova Scotia (BNS), Imperial Oil Limited (IMO), Sun
Life Financial Inc. (SLF) and TC Energy Corporation (TRP).
These companies are included in the S&P/TSX 60 Index of
Toronto Stock Exchange. Each set has data of 1416 days
(from December 31, 2013 to August 21, 2019), presenting
the closing price for each day. For method implementation,
1400 data (from December 31, 2013 to July 29, 2019) were
separated for training purposes. The remaining 16 data (from
July 30, 2019 to August 21, 2019) were kept for testing pur-
poses. Furthermore, four subsets from testing data were taken
at four future horizons, i.e., h = 4, 8, 12 and 16. Training data
of all five sets are graphically presented in Fig. 1. These data
are also concisely presented in Table 1 with summary statis-
tics, including mean, median, minimum, maximum, Coeffi-
cient of Variation (or briefly COV), skewness, and kurtosis.
Furthermore, typical EMD components (for BMO data set)
are shown in Fig. 2, reflecting the price dynamics from a
granular point of view.
Fig.1 portrays that all the series are fluctuating oscillations
of upward and downward. It can also be seen that all the
stock price effect by mini-recession on 2015-2016 before
moving upward again. Meanwhile, Table 1 shows that the
mean of BMO closing stock prices is the highest, but all
stocks have similar COV around 0.1, which indicate a similar
investment risk. Besides, the five stocks’ distribution are not
FIGURE 1. Closing price graphs for five stocks (a) BMO, (b) BNS, (c) IMO,
symmetric, where some stocks are positively skewed, and (d) SLF and (e) TRP.
some are negatively skewed. Moreover, the distributions are
also platykurtic. Also, Fig. 2 displays EMD decomposed
components of the BMO data set from high frequency to low
two experimental methods, only EMD-ARIMA-EWMA out-
frequency with the residual to capture the data’s remaining
performed all other methods considering all error values and
noise and essential features.
all datasets. In Fig. 3, forecasted data and test data are pre-
sented for all six methods and five datasets where ARIMA
B. RESULTS and EWMA results are very close which seem overlapped in
Table 2 and Table 3 present the empirical results as found the graph. Table 4 presents computation time for each dataset
in this study. These tables contain forecast accuracy tools for h = 16 for each method (along with EMD algorithm).
RMSE and MAE for measuring absolute errors and MAPE The computation time of EWMA was the lowest among all
as relative error measure. A small error value indicates bet- methods and the highest time required for EMD-ARIMA.
ter accuracy or better performance. They also contain five Hybrid methods require more time than single methods. In the
data sets (BMO, BNS, IMO, SLF and TRP) and six meth- following section, performance accuracies and overall results
ods (i.e., two benchmark methods – ARIMA and EWMA; regarding methods and datasets are discussed, along with
two EMD-based hybrid methods EMD-ARIMA and EMD- computation time (in second).
EWMA; two experimental methods EMD-ARIMA-EWMA
and EMD-EWMA-ARIMA). Table 2 presents the methods’ V. DISCUSSION AND OUTCOME OF THE STUDY
performance results at horizon h = 4 and 8, while Table 3 Experimental results (Table 2 and Table 3) found for five
contains results related to horizon h = 12 and 16. Between sets of stock price data reveal the superiority of proposed
VOLUME 9, 2021 132323
M. R. Hossain et al.: Improving Stock Price Prediction Using Combining Forecasts Methods
TABLE 1. Summary statistics for five daily closing stock prices.
TABLE 2. Forecasting accuracy of five stock data sets for forecast horizon h = 4 and h = 8.
EMD-ARIMA-EWMA among all the six methods in all the vary. For first dataset BMO, the best performing EMD-
forecast horizons h = 4, 8, 12 and 16. The smaller the ARIMA-EWMA produced relative MAPE errors 1.194,
errors (RMSE and MAE for absolute; MAPE for relative 2.428, 3.590 and 4.256 at h = 4, 8, 12 and 16 respec-
measure), the better the method. Statistical methods effec- tively. The second-best performing method is EWMA with
tively synergized with EMD components, i.e., ARIMA and MAPE 2.158, 3.647, 4.834 and 5.502 at the same h val-
EWMA contributed better performance in high-frequency ues in the same order. Furthermore, for other datasets (i.e.,
IMFs and low-frequency IMFs along with residual respec- BNS, IMO, SLF and TRP), EMD-ARIMA-EWMA has the
tively as expected. smallest error result of RMSE, MAPE, or MAE value all h
Each dataset is intrinsically different, and therefore fore- locations. However, at all h values, the second-best performer
casting performance of different methods on them usually was EMD-ARIMA for BNS data. Also, for IMO and SLF
132324 VOLUME 9, 2021
M. R. Hossain et al.: Improving Stock Price Prediction Using Combining Forecasts Methods
TABLE 3. Forecasting accuracy of five stock data sets for forecast horizon h = 12 and h = 16.
TABLE 4. Computation time (in second) of six methods (and EMD algorithm) for five stock data sets for forecast horizon h = 16.
datasets, EMD-ARIMA performed as a second-best method All datasets are not equally forecastable due to their char-
at all h values. Finally, for TRP dataset considering MAPE acteristic features. Forecastability rank or order can be found
results, EMD-ARIMA was second-best performing method by comparing the least MAPE value for each data at each
at h = 4, 8 and 12; but EMD-EWMA was a second-best horizon. Therefore, with the MAPE results of the outper-
performer at h = 16. From Table 2-3, on average, the worst- forming method EMD-ARIMA-EWMA, the best-forecasted
performing method was EMD-EWMA-ARIMA, except for datasets at different forecast horizons are TRP at horizon
SLF data at h = 4 and TRP data at h = 4 and 8. h = 4 (MAPE 0.395), BNS at other three horizons h = 8
VOLUME 9, 2021 132325
M. R. Hossain et al.: Improving Stock Price Prediction Using Combining Forecasts Methods
FIGURE 2. EMD component (IMFs and residual) graphs for Bank of
Montreal (BMO).
(MAPE 0.656), h = 12 (MAPE 0.702) and h = 16 (MAPE
0.699). Similarly, least forecasted stock datasets with the best
performing EMD-ARIMA-EWMA method are SLF at h = 4 FIGURE 3. Forecast data graphs of six methods for five datasets.
(MAPE 1.970), IMO at other horizons h = 8 (MAPE 4.220),
h = 12 (MAPE 5.114) and h = 16 (MAPE 4.601).
In any hybrid method, efficient constituent methods can a single method (here ARIMA or EWMA) is used for
improve the overall performance, and best combination can each EMD component which contributes to much computa-
produce the best results. Likewise, in combining EMD tion time. In their hybrid method, [24] found that EWT-Q-
with ARIMA and EWMA, EMD-ARIMA-EWMA is poten- BPNN (where EWT stands for Empirical Wavelet Transform,
tially best suited to perform better. However, EMD-EWMA- Q stands for Q-learning algorithm and BPNN stands for Back
ARIMA is ill-suited and perform poorly. This was an Propagation Neural Network) method required 614.63s while
essential part of our experimental investigation. ARIMA required 13.55s for their Series #1 dataset. Here, for
It should be expected that due to their additional BMO dataset, the proposed EMD-ARIMA-EWMA required
computation or computational complexity, the hybrid or com- 1.96s while ARIMA required 0.106s. However, EMD algo-
bination methods require more computation time in compar- rithm required only 0.005s. Therefore, constituent ARIMA
ison to single methods. In an EMD-based hybrid method, method contributed more time to the proposed method. Thus,
132326 VOLUME 9, 2021
M. R. Hossain et al.: Improving Stock Price Prediction Using Combining Forecasts Methods
the proposed or other EMD-based similar methods can fur- [6] H. Liu, C. Li, Y. Shao, X. Zhang, Z. Zhai, X. Wang, X. Qi, J. Wang,
ther be improved using a suitable and effective combination Y. Hao, Q. Wu, and M. Jiao, ‘‘Forecast of the trend in incidence of acute
hemorrhagic conjunctivitis in China from 2011–2019 using the seasonal
of parametric, non-parametric and semi-parametric methods autoregressive integrated moving average (SARIMA) and exponential
(e.g., as presented by [25]) with EMD components. However, smoothing (ETS) models,’’ J. Infection Public Health, vol. 13, no. 2,
this study aimed to find an EMD-based better forecasting pp. 287–294, Feb. 2020.
[7] Q. T. Tran, L. Hao, and Q. K. Trinh, ‘‘A comprehensive research on
method which approach can further be extended to develop exponential smoothing methods in modeling and forecasting cellular traf-
hybrid methods that will be optimized with forecast accuracy fic,’’ Concurrency Comput., Pract. Exper., vol. 32, no. 23, Dec. 2020,
and computation time. Art. no. e5602.
[8] S.-M. Chen, X.-Y. Zou, and G. C. Gunawan, ‘‘Fuzzy time series forecasting
Complementing overall discussion with the empiri- based on proportions of intervals and particle swarm optimization tech-
cal results, the study’s outcome is that EMD-ARIMA- niques,’’ Inf. Sci., vol. 500, pp. 127–139, Oct. 2019.
EWMA method had comparatively high potentiality over [9] J. W. Koo, S. W. Wong, G. Selvachandran, H. V. Long, and L. H. Son,
other methods to produce better forecast accuracy. Rele- ‘‘Prediction of air pollution index in Kuala Lumpur using fuzzy time series
and statistical models,’’ Air Qual., Atmos. Health, vol. 13, no. 1, pp. 77–88,
vantly, we assume that adaptive decomposition EMD-based Jan. 2020.
hybridization with different well-suited statistical methods [10] F. M. Khan and R. Gupta, ‘‘ARIMA and NAR based prediction model for
can forecast better if these methods are applied considering time series analysis of COVID-19 cases in India,’’ J. Saf. Sci. Resilience,
vol. 1, no. 1, pp. 12–18, Sep. 2020.
frequencies of the EMD components. Furthermore, although [11] Ü. Ç. Büyükşahin and Ş. Ertekin, ‘‘Improving forecasting accuracy
the proposed method is employed on stock market data, this of time series data using a new ARIMA-ANN hybrid method and
or other potentially similar methods can be practicable in empirical mode decomposition,’’ Neurocomputing, vol. 361, pp. 151–163,
Oct. 2019.
other forecasting related fields of time series application.
[12] J. Wang and J. Wang, ‘‘Forecasting stochastic neural network based on
financial empirical mode decomposition,’’ Neural Netw., vol. 90, pp. 8–20,
VI. CONCLUDING REMARKS Jun. 2017.
[13] Z.-X. Wang, Y.-F. Zhao, and L.-Y. He, ‘‘Forecasting the monthly iron ore
Time series forecasting is challenging, especially when it
import of China using a model combining empirical mode decomposition,
holds excessive uncertainty and contains less capturable pat- non-linear autoregressive neural network, and autoregressive integrated
terns. However, in many cases, efficient methods can capture moving average,’’ Appl. Soft Comput., vol. 94, Sep. 2020, Art. no. 106475.
dominating patterns and features. The local adaptation fea- [14] J. Cao, Z. Li, and J. Li, ‘‘Financial time series forecasting model based
on CEEMDAN and LSTM,’’ Phys. A, Stat. Mech. Appl., vol. 519,
ture of the EMD algorithm has effective pattern extracting pp. 127–139, Apr. 2019.
capability, which can play an essential role in serving or [15] N. Nava, T. Matteo, and T. Aste, ‘‘Financial time series forecasting using
boosting other methods’ efficiency. This study found that empirical mode decomposition and support vector regression,’’ Risks,
vol. 6, no. 1, p. 7, Feb. 2018.
experimented method EMD-ARIMA-EWMA outperformed
[16] S. Makridakis, E. Spiliotis, and V. Assimakopoulos, ‘‘The m4 competition:
the results of single ARIMA, EWMA, hybrid EMD-ARIMA Results, findings, conclusion and way forward,’’ Int. J. Forecasting, vol. 34,
and EMD-EWMA. This reflection was found through the no. 4, pp. 802–808, Oct. 2018, doi: 10.1016/[Link].2018.06.001.
results related to all the five data sets and four forecast hori- [17] S. Makridakis, E. Spiliotis, and V. Assimakopoulos, ‘‘Statistical and
machine learning forecasting methods: Concerns and ways forward,’’ PLoS
zons. Hopefully, this research study will inspire further EMD- ONE, vol. 13, no. 3, Mar. 2018, Art. no. e0194889.
based effective method development and will get the attention [18] N. E. Huang, Z. Shen, S. R. Long, M. C. Wu, H. H. Shih, Q. Zheng,
of researchers and practitioners of time series from different N. C. Yen, C. C. Tung, and H. H. Liu, ‘‘The empirical mode decomposition
and the Hilbert spectrum for nonlinear and non-stationary time series
fields beyond the financial or economic domain. In the future, analysis,’’ Proc. Roy. Soc. London Ser. A, Math., Phys. Eng. Sci., vol. 454,
we plan to investigate more on EMD-based methods or their no. 1971, pp. 903–995, Mar. 1998, doi: 10.1098/rspa.1998.0193.
existing problems yet unexplored and contribute to analysis [19] N. E. Huang, M.-L. Wu, W. Qu, S. R. Long, and S. S. P. Shen, ‘‘Appli-
cations of Hilbert–Huang transform to non-stationary financial time series
and forecasting of time series empirical data as well as in
analysis,’’ Appl. Stochastic Models Bus. Ind., vol. 19, no. 3, pp. 245–268,
geophysical events. Jul. 2003, doi: 10.1002/asmb.501.
[20] Y. Fang, B. Guan, S. Wu, and S. Heravi, ‘‘Optimal forecast combina-
tion based on ensemble empirical mode decomposition for agricultural
REFERENCES commodity futures prices,’’ J. Forecasting, vol. 39, no. 6, pp. 877–886,
[1] T. Alghamdi, K. Elgazzar, M. Bayoumi, T. Sharaf, and S. Shah, ‘‘Fore- Sep. 2020.
casting traffic congestion using ARIMA modeling,’’ in Proc. 15th [21] A. M. Awajan, M. T. Ismail, and S. A. Wadi, ‘‘Improving forecasting
Int. Wireless Commun. Mobile Comput. Conf. (IWCMC), Jun. 2019, accuracy for stock market data using EMD-HW bagging,’’ PLoS ONE,
pp. 1227–1232. vol. 13, no. 7, Jul. 2018, Art. no. e0199582.
[2] M. Alsharif, M. Younes, and J. Kim, ‘‘Time series ARIMA model for [22] S. Abadan and A. Shabri, ‘‘Hybrid empirical mode decomposition-
prediction of daily and monthly average global solar radiation: The ARIMA for forecasting price of rice,’’ Appl. Math. Sci., vol. 8,
case study of Seoul, South Korea,’’ Symmetry, vol. 11, no. 2, p. 240, pp. 3133–3143, 2014, doi: 10.12988/ams.2014.43189.
Feb. 2019. [23] M. R. Hossain and M. T. Ismail, ‘‘Empirical mode decomposition based on
[3] A. Hernandez-Matamoros, H. Fujita, T. Hayashi, and H. Perez-Meana, theta method for forecasting daily stock price,’’ J. Inf. Commun. Technol.,
‘‘Forecasting of COVID19 per regions using ARIMA models and vol. 19, no. 4, pp. 533–558, Aug. 2020.
polynomial functions,’’ Appl. Soft Comput., vol. 96, Nov. 2020, [24] H. Liu, C. Yu, C. Yu, C. Chen, and H. Wu, ‘‘A novel axle temperature
Art. no. 106610. forecasting method based on decomposition, reinforcement learning opti-
[4] J. W. Miller, ‘‘ARIMA time series models for full truckload transportation mization and neural network,’’ Adv. Eng. Informat., vol. 44, Apr. 2020,
prices,’’ Forecasting, vol. 1, no. 1, pp. 121–134, Sep. 2018. Art. no. 101089.
[5] J. Köppelová and A. Jindrová, ‘‘Application of exponential smoothing [25] J. Li, X. Xia, W. K. Wong, and D. Nott, ‘‘Varying-coefficient semiparamet-
models and arima models in time series analysis from Telco area,’’ Agris ric model averaging prediction,’’ Biometrics, vol. 74, no. 4, pp. 1417–1426,
Line Papers Econ. Informat., vol. 11, no. 3, pp. 73–84, Sep. 2019. Dec. 2018.
VOLUME 9, 2021 132327
M. R. Hossain et al.: Improving Stock Price Prediction Using Combining Forecasts Methods
MOHAMMAD RAQUIBUL HOSSAIN received SAMSUL ARIFFIN BIN ABDUL KARIM is cur-
the B.S. and M.S. degrees from the Department rently a Senior Lecturer with Universiti Teknologi
of Mathematics, University of Dhaka, Bangladesh, PETRONAS, Malaysia. He is a Professional
in 2006 and 2007, respectively. He is currently pur- Technologist registered with Malaysia Board
suing the Ph.D. degree with the School of Mathe- of Technologists (MBOT). He is a Certified
matical Sciences, Universiti Sains Malaysia, Pulau WOLFRAM Technology Associate, Mathemat-
Pinang, Malaysia. His research interests include ica Student Level. He has published more than
time series forecasting, predictive data analytics, 140 papers in journals and conferences, includ-
data science, and machine learning. ing three edited conferences volume and 60 book
chapters. He has also published ten books with
springer publishing, including five books with studies in systems, decision
and control (SSDC) series, and one book with UTP Press. He was a recipient
of the Effective Education Delivery Award and the Publication Award (jour-
nal and conference paper), UTP Quality Day 2010, 2011, and 2012.
MOHD TAHIR ISMAIL is currently an Associate
Professor and a Researcher with the School of
Mathematical Sciences, Universiti Sains Malaysia.
He has published more than 150 publications in
reviewed journals and proceedings (some of them
are listed in ISI, Scopus, Zentralblatt, MathSciNet,
and other indices). His research interests include
financial time series, econometrics, categorical
data analysis, and applied statistics. He is currently
an Exco Member of the Malaysian Mathematical
Sciences Society and an active member of other scientific professional
bodies.
132328 VOLUME 9, 2021