0% found this document useful (0 votes)
57 views9 pages

Stock Price Prediction with Bi-LSTM

This document proposes using bi-directional LSTM based sequence to sequence modeling and multitask learning to predict stock prices. It introduces stock markets and the challenges of predicting stock prices, which depend on many static and dynamic factors. It then proposes two systems: 1) A bi-directional LSTM seq2seq model to map input sequences of stock price data to output sequences to predict future stock prices. 2) A multitask learning model to determine relationships between different stock price attributes (open, high, low, close) and predict them simultaneously. The models are tested on Indian stock market data and outperform other algorithms in predicting prices.

Uploaded by

Rishabh Khetan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
57 views9 pages

Stock Price Prediction with Bi-LSTM

This document proposes using bi-directional LSTM based sequence to sequence modeling and multitask learning to predict stock prices. It introduces stock markets and the challenges of predicting stock prices, which depend on many static and dynamic factors. It then proposes two systems: 1) A bi-directional LSTM seq2seq model to map input sequences of stock price data to output sequences to predict future stock prices. 2) A multitask learning model to determine relationships between different stock price attributes (open, high, low, close) and predict them simultaneously. The models are tested on Indian stock market data and outperform other algorithms in predicting prices.

Uploaded by

Rishabh Khetan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Stock Price Prediction using Bi-Directional LSTM

based Sequence to Sequence Modeling and Multitask


Learning
Siddartha Mootha Sashank Sridhar Rahul Seetharaman Dr. S. Chitrakala
Department of Computer Science Department of Computer Science Department of Computer Science Department of Computer Science
and Engineering and Engineering and Engineering and Engineering
College of Engineering Guindy, College of Engineering Guindy, College of Engineering Guindy, College of Engineering Guindy,
Anna University Anna University Anna University Anna University
Chennai, India Chennai, India Chennai, India Chennai, India
siddartha.mootha20@[Link] [Link]@[Link] rahulseetharaman@[Link] [Link]@[Link]

Abstract— The stock market is a dynamic and volatile The stock market is not static, it is an extremely dynamic
platform which provides an environment for traders to invest and environment [8]. The market is non-linear and can be highly
trade in shares. The price of a stock is dependent on numerous unpredictable [9]. The financial market provides an environment
static and dynamic features. Predicting the future price of a for traders to invest and trade in shares, stocks, securities and
particular company’s stock can be extremely beneficial for equities. The stock market provides an environment where a
traders. Seq2Seq modelling helps map an input sequence to an small investment could lead to a large return, or even a large
output sequence. In this paper, we propose a system to predict the loss. Predicting the future price [10] of a particular company’s
future Open, High, Close, Low (OHCL) value of a stock using a stock could be extremely beneficial for investors, nonetheless it
Bi-Directional LSTM based Sequence to Sequence Modelling.
is extremely challenging. The price of a stock varies on
Each OHCL price is an independent sequence and multitask
numerable aspects [11] including but not limited to weather
learning helps map the interrelations between them. A multitask
system is also proposed which uses sub tasks and shared tasks to forecasts, speech by a politician, social media, news articles,
model the prices. Stock prices of Tata Consumer Products global economy etc. As can be seen, most of the above
Limited from the National Stock Exchange (NSE) of India is used. mentioned attributes are not controllable by an average human.
To evaluate the efficiency of the proposed systems, they are Nonetheless, with the advancement in artificial intelligence and
compared against various machine learning algorithms. The machine learning, and the availability of high performance super
proposed Seq2Seq and multitask systems comfortably outperform computers, predicting the future price of a stock is possible [12].
the existing algorithms with RMSE values of 3.98 and 7.87
A general stock has 4 primary price features [13], i.e. Open
respectively.
Price, High Price, Low Price, Close Price (OHLC). Based on
Keywords—Stock Forecasting, Sequence to Sequence these 4 primary price attributes, this paper proposes various
Modelling, Multitask Learning, Bi-directional LSTM, Prediction systems to predict the future price of a stock.
System Sequence to Sequence (Seq2Seq) [14] modelling helps map
I. INTRODUCTION an input sequence of data to an output sequence of data. The
Seq2Seq models use an encoder-decoder framework [15]. The
The stock market is an extremely potent environment encoder helps map the structure of the input sequence and the
capable of handling millions of transactions in a short span of decoder generates the corresponding output sequence. The
time [1]. Every transaction is routed through a central exchange encoder and decoder make use of Recurrent Neural Networks
[2]. The stock market involves three primary parties i.e. a buyer, (RNN) [16] for analysing the sequence of the data. Bi-
seller and an exchange [3]. Generally a buyer purchases a stock directional LSTMs (Bi-LSTM) [17] are a form of RNNs that
via a third party agent such as a broker [4]. A stock is a measure map the sequences in both forward and reverse directions. This
of ownership of a company [5]. helps the network deal with increased sequence length and long
There are different types of companies [6] such as statutory term dependencies.
companies, Royal Chartered Companies, Private Limited A Bi-directional LSTM (Bi-LSTM) model which maps the
Companies, Public Limited Companies etc. When a company input sequence of OHLC prices for a particular day is proposed
decides to offer shares of stock to the general public, it must in this paper. The proposed Seq2Seq model aims to generate an
launch an Initial Public Offering (IPO) [7]. This offering states output sequence of OHLC data for the day under study. In stock
the number of shares available and the price at which it is price prediction, the future price is as important as the past price.
available. Although the price of a stock is extremely volatile, it By creating a window based model [18], the past and future
can change drastically from time to time. prices are modelled using Bi-LSTMs.

978-1-7281-9656-5/20/$31.00 ©2020 IEEE


0078

Authorized licensed use limited to: ANNA UNIVERSITY. Downloaded on December 26,2020 at 05:22:48 UTC from IEEE Xplore. Restrictions apply.
Multitask learning [19] is a machine learning paradigm that the benchmark for the prediction. The pre-processed data was
works effectively if the parameter space consists of parameters put into a one way and two way LSTM. Dropout is an
which are interrelated. It is similar to how humans learn optimization technique to prevent overfitting. Dropout
something. For example, when humans learn to ride a car, they technique was applied on the neural network. The proposed
understand how to use gears, clutches, accelerators, etc. This BLSTM was compared against the ARIMA model as well as a
knowledge can help them to learn to ride a motorbike as well. It traditional Long Short Term Memory (LSTM). To measure the
can improve performance by learning to solve overlapping performance of the proposed system, RMSE, MAE and
subproblems and using the knowledge obtained in solving one deviation accuracy was calculated. It could be seen from the
sub-problem to effectively use it for another sub-problem, in a results that the proposed system outperformed the ARIMA
joint learning process [20]. model as well as the LSTM model. RMSE and MAE were
reduced by 24.2% and 19.4% respectively.
In this paper, the multitask model is established to determine
the variation of OHLC prices. By determining the relative K. A. Althelaya et al. [17] proposed a Bidirectional LSTM
changes in OHLC prices, the prediction error can be reduced. for Short- and Long-Term Stock Market Prediction. The authors
Multitask learning helps identify the interrelation between the had made use of the Standard and Poor 500 Index (S&P500)
OHLC prices by predicting the OHLC prices simultaneously. As historical data for their proposed work. The S&P500 is a leading
the prediction of each of the OHLC prices is viewed as an indicator for various top traded companies. The dataset was
independent task and as the OHLC prices are dependent on one normalised and pre-processed. At the end of every trading day,
another, this forms a multitask learning problem. there is a closing price. The closing price was used for the sake
of evaluating the performance of the system. The authors
To test the performance of the system, the two models are proposed two systems, one for short term prediction and one for
compared against various algorithms such as k-NN, Bayesian long term prediction. To evaluate the performance of both the
Regression, Ridge Regression, Linear Regression, Huber models, it was evaluated and compared against 3 different
Regression, TheilSen Regression, Lasso Regression, Least models, i.e. MLP-ANN, LSTM and SLSTM. The BLSTM
Angle Regression and SVM. The results of the comparison show achieved a lower RMSE and MAE score in comparison to the
how well the proposed models perform by outperforming the other models. The BLSTM outperformed the existing models
basic Machine Learning models. and achieved better convergence for short term and long term
The rest of the paper is organized as follows: In Section II, predictions.
we will discuss the related works carried out in stock price Sequence to sequence modelling helps the network map a
prediction. Section III explains about the design of the system in sequential input to a corresponding sequential output. C.-H. Cho
detail. Section IV talks about the implementation details. et al. [23] modelled the Taiwan stock price information of 932
Section V analyses the results. In Section VI, the conclusion and companies in the period between 2007 and 2019. They trained
future works are discussed. three different models : LSTM, Seq2Seq and WaveNet using
II. RELATED WORKS daily trading data as well as technical indicators. They found the
correlation between the predicted price and actual price and used
Stock price prediction can be done using Artificial Neural the correlation as a performance metric. The WaveNet model
Networks (ANN). ANNs use adaptive weights to forecast stock proposed by them outperformed the other two models.
prices. Y. Bing et al. [21] proposed an ANN to predict the index
of the Shanghai Stock Exchange. The authors studied the market Every stock exchange has an index. The market index is an
between March 17, 2010 to April 28, 2010. They considered 5 extremely critical factor for investors. J. Eapen et al. [24]
features of the market, open, high, close, low and volume. The proposed a deep learning model with Convolutional Neural
neural network constructed was successful in predicting the Networks (CNN) and Bi-Directional LSTM (Bi-LSTM) to
daily lowest, highest, and closing value of the Shanghai Stock predict the future index of the market. The proposed model was
Exchange. The gradient search technique and learning algorithm implemented on the S&P 500 grand challenge dataset. Multiple
were constructed in the deep learning model. The results showed pipelines of CNN and BLSTM units were combined. The
that the model was successful in predicting the daily low, high, authors presented various combinations of single and multiple
and closing value of the Shanghai Stock Exchange for a short pipeline deep learning models applied on various kernel sizes
time period, but inefficient in predicting the return rate of the for CNN and for varying numbers of Bi-LSTM units. It was
Shanghai Stock Exchange. observed that the deep network created due to CNNs, when
combined with Bi-LSTM created a more efficient system than
RNNs help identify the relationship present in sequential previously existing ones.
data. Bi-LSTMs are a subset of RNNs which model the data in
forward and reverse directions. This takes into account what the As shown by J. Eapen et al. [24], Bi-LSTMs increase the
past stock prices were and what the future stock prices might be. accuracy of prediction by mapping the sequential data in both
M. Jia et al. [22] proposed a framework which made use of the directions. In our first model, we have set up a Seq2Seq model
bidirectional long-short term memory (BLSTM) neural network which uses Bi-LSTMs to map the input sequential data.
for predicting the future price of a stock. The authors used the C. Li et al. [25] built a multi-task RNN to extract features
historical data of the GREE stock. They collected data for 568 from OHLC, volume and amount data of a particular stock.
days from January 1, 2017 to May 14, 2019. The data consisted Features extracted from each individual task were then
of 14 features such as open, high, close, volume etc. The data concatenated and mapped to predict future prices. The multi-
was normalized and pre-processed. The close value was used as task model was a “MultiTask Market Price Learner” and

0079

Authorized licensed use limited to: ANNA UNIVERSITY. Downloaded on December 26,2020 at 05:22:48 UTC from IEEE Xplore. Restrictions apply.
Fig. 1. Bi-Directional LSTM based Encoder-Decoder

Fig. 2. Multitask Learning Model

consisted of three dual stage attention based recurrent neural It can be seen from C. Li et al. [25] that multitask learning
networks (DARNN). DARNNs help extract temporal features helps model individual tasks and the overall interrelation
and were shown to perform better than LSTMs and attention between the tasks. In our second model, we use multitask
based LSTMs. Two types of tasks were handled by the model learning to model the OHLC data as individual tasks and to
proposed by the authors. The low-level tasks were used to predict the future OHLC values. As we use multitask learning
predict volatility and stock prices in the future. The higher level the variations in OHLC prices are interdependent on each other.
task made use of the low-level tasks to predict the movement of
stock prices. The model was trained on data from the CSI 200, The price of a stock can vary based on factors such as blogs,
300 and 500 indices of the Chinese stock market. The Technical news, articles, rumours etc. The numerous external factors make
Analysis Library in python was used to extract 8 technical it extremely difficult to predict the price of a stock. W. Khan et
indicators for each stock. These features along with the market al. [26] proposed a system to predict the market index as well as
features i.e. OHLC, Volume and Amount, were trained using the the price of a stock based on social media news. Impact of social
multi-task learner. The multi-task learning model achieved an media and financial news on a stock is studied. Deep learning
accuracy of 63.67% for trends and 62.50% for volatility. and an ensemble of classifiers is used for prediction. The results
show that social media influences the New York and IBM stocks

0080

Authorized licensed use limited to: ANNA UNIVERSITY. Downloaded on December 26,2020 at 05:22:48 UTC from IEEE Xplore. Restrictions apply.
the most, whereas financial news influences the London and decoded by the LSTM units of the decoder. The decoder tries to
Microsoft stocks the most. predict a single sequence of OHLC prices.
It is observed from W. Khan et al. [26] that stock prices are D. Multitask Learning
affected by multiple factors. By using multitask learning, we try Multitask learning involves learning sub tasks and training a
to model the impact of each of the OHLC prices on the overall common shared task using the information learnt [32]. In the
stock price, individually. The lower level tasks in our model implemented system as seen in Figure 2, each of the OHLC
maps the individual OHLC prices and the higher level task prices is taken as a subtask. Input sequence is prepared using a
forecasts the OHLC prices. sliding window of a particular window length. The window
III. DESIGN METHODOLOGY length is chosen such that the loss is minimal. Sequences of
length n for each price are trained with the output being the
A. Dataset Description + 1 price in the sequence. Each sub task is modelled using
Stock prices for every company listed on a stock exchange a Bi-LSTM model. The outputs of the individual sub tasks are
is readily available. This paper makes use of the stock prices of concatenated and given to a shared Bi-LSTM model which
Tata Consumer Products Limited from the National Stock predicts the output sequence for each of the OHLC prices. In this
Exchange (NSE) of India [27]. The stock prices were extracted way the features learnt from each sub task is propagated to a
in the period between 2013 and 2018. The dataset contains four common shared task which is responsible for determining the
main features - Open, High, Low and Close Price (OHLC) as output.
seen in Table I. IV. IMPLEMENTATION
TABLE I. DATA ATTRIBUTES Implementation of the proposed models was done on a
Windows 10 machine with Intel I7 8th Gen Processor, 16 GB
S. No. Feature Description
Opening price of the stock on a particular RAM and NVidia GEForce GTX 1060 graphics processor.
1 Open Price
day
A. Preprocessing the Dataset
Close The closing price of the stock at the end
2
Price of a trading day The dataset collected has OHLC prices for a total of 1150
3 High Price The peak price on a particular day samples. The dataset was split into a training and testing set with
4 Low Price
The lowest price of the stock on a a ratio of 75:25. The dataset is cleaned to remove any null
particular day values. The data is then normalized using the MinMaxScaler of
sklearn [33]. This scaling subtracts the minimum value in the
B. Sequence to Sequence Modelling dataset from every value in the dataset, and divides every value
in the dataset by the maximum value in the dataset. This
Sequence to sequence modelling tries to forecast the procedure was repeated across multiple time series, each of
consecutive output of the given input sequence [28]. Seq2Seq which denoted Opening Price, Closing Price, High Price, and
modelling tries to predict the output sequence using a given Low Price.
input sequence [29]. The problem solved by Seq2Seq mapping
is that sequences of variable length can be mapped to the output B. Determining the Sliding Window
sequence. Seq2Seq models can be used in a variety of tasks, A sequence of OHLC prices was created by using a sliding
mainly those that require generation of text, sequences, etc [30]. window. For all the values at every time step, a window size of
C. Bi-Directional LSTM based Encoder-Decoder was chosen as the input and the next value after it, + 1 ,
was chosen as the output. This procedure was repeated as part
Encoder-Decoder [31] is a method to solve Seq2Seq
of a sliding window procedure and the input and output was
predictions. The encoder converts a variable length input
generated, which was in turn fed to both the constructed models.
sequence to a fixed-length vector which is decoded by the
decoder. The decoder tries to predict the output sequence by
using the encoded input. This method helps map variable length
sequences.
Figure 1 shows the overall architecture diagram of the
implemented Seq2Seq model. OHLC prices of stocks for each
day form a sequence. An input sequence is taken as a sequence
of OHLC prices for a particular window length of days. The
output sequence is the sequence of OHLC prices for the next day
after the input window.
The encoder consists of Bi-LSTM units, each of which maps
the Open, Close, High and Low price. The Bi-LSTM units try to
model the relation between the past sequences and the future
sequences. At a particular instant, a sequence of window length
units are given to the encoder. The window length is chosen such Fig. 3. Variation of RMSE with Window Size in Bi-LSTM based Seq2Seq
that the output prediction error is minimal. The encoder converts Model
the input sequence to a fixed length vector. This vector is

0081

Authorized licensed use limited to: ANNA UNIVERSITY. Downloaded on December 26,2020 at 05:22:48 UTC from IEEE Xplore. Restrictions apply.
From Figure 3, it can be seen that the RMSE is lowest at a
window size of 60 for the Bi-LSTM based Sequence to
Sequence model.
Algorithm 1 shows how the dataset is created for the
Seq2Seq model. The training set is taken as a sequence of
determined window size. The training outputs have all 4 of the
OHLC prices together.
Algorithm 1 : Dataset for Seq2Seq model Fig. 5. Implementation of Seq2Seq model

From Figure 4, it can be seen that the RMSE is lowest at a


window size of 60 for the Bi-LSTM based Multitask model.

Fig. 6. Variation of RMSE with Number of Neurons in the Encoder

The encoder returns two vectors [34] of fixed size. One is a


hidden state vector and another is a context vector. The decoder
is a LSTM layer with 512 nodes. The decoder takes the output
states of the encoder and returns a sequence of output values. A
dense layer is used to map the decoder outputs to the OHLC
values of the output sequence. The Seq2Seq model is trained
Fig. 4. Variation of RMSE with Window Size in Bi-LSTM based Multitask
for 10 epochs.
Model D. Bi-LSTM based Multitask Model
Algorithm 2 : Dataset for Multitask model Figure 7 shows the overall implementation of the multitask
model. In the multitask model, each sub task is a Bi-LSTM
model that tries to learn each of the OHLC prices. Each sub task
has 2 Bi-LSTM layers with 128 and 64 nodes respectively.
From Figure 8 it can be seen that the RMSE is lowest at 2
hidden layers. Each Bi-LSTM layer is followed by a dropout of
0.2.
The final output from the 4 sub tasks is a 4 * 1 dimensional
vector that predicts open, close, high and low prices. This
output vector is fed to a shared task that tries to reduce the loss
of predictions further. This shared task has 2 Dense layers with
Algorithm 2 shows how the dataset is created for the 100 and 50 nodes each. It can be seen from Figure 9 that the
Multitask model. The training set is taken as a sequence of RMSE is lowest when the shared task has 2 hidden [Link]
determined window size. The training outputs are segregated shared task is mapped to the OHLC prices using 4 output Dense
separately for OHLC prices. This is because prediction of each layers as seen in Figure [Link] multitask model is trained for 5
price forms a sub task in the multitask model. epochs.
C. Bi-LSTM based Seq2Seq Model E. Parameters for Training
Figure 5 shows the implementation of the Seq2Seq model. Table II shows the parameters used for training the models.
The Bi-LSTM based Seq2Seq model makes use of an encoder- The optimizer used is RMSProp [35]. RMSProp maintains a
decoder architecture. The encoder is Bi-directional LSTM layer moving average of the square of gradient values and divides the
with 256 nodes. Figure 6 shows the variation of RMSE with the gradient by the root of the moving average. A learning rate of
number of neurons in the encoder. It can be seen that the RMSE 0.001 was used for training the model over the generated data.
is lowest when the number of neurons in the encoder is 256. The loss function used was the mean squared error loss function

0082

Authorized licensed use limited to: ANNA UNIVERSITY. Downloaded on December 26,2020 at 05:22:48 UTC from IEEE Xplore. Restrictions apply.
[36], which computes the mean of the squared difference V. RESULTS AND ANALYSIS
between the label and predictions. The squared difference is
done to avoid positive and negative loss differences being
nullified across the dataset.

Fig. 7. Implementation of Multitask model

Fig. 8. Variation of RMSE with Number of Hidden Layers for Sub Tasks

Fig. 10. Prediction curves for a) Open b) Close c) High and d) Low prices by
Bi-LSTM based Seq2Seq model

Figures 10 and 11 show the prediction curves of OHLC data


Fig. 9. Variation of RMSE with Number of Hidden Layers in Shared Task for both the models. The green line shows the actual data and
the orange line shows the predicted data. The closeness of the
TABLE II. PARAMETERS FOR TRAINING two lines determines how well the models predict the outputs.
S. No. Parameter Value For any machine learning problem, it is rudimentary and critical
1 Optimizer RMSProp to evaluate the performance of the proposed model. The
2 Loss Function Mean Squared Error proposed system is evaluated using the following metrics where
3 Learning Rate 0.001 ‘n’ represents the number of observations, represents the
4 Batch Size 32

0083

Authorized licensed use limited to: ANNA UNIVERSITY. Downloaded on December 26,2020 at 05:22:48 UTC from IEEE Xplore. Restrictions apply.
predicted price of the stock, and represents the actual value The proposed seq2seq and Bi-LSTM multitask model is
of the stock. compared against existing regression models. In Machine
Learning, regression [38] is a paradigm where future values are
predicted by studying and learning the trends of the existing
values. The proposed system is compared against regression
models such as k-NN Regressor, Bayesian Ridge Regressor,
Ridge Regression, Linear Regression, Huber Regression,
TheilSen Regressor, Random Sample Consensus, Orthogonal
Matching Pursuit, Lasso Regression, SVM Regression, Passive
Aggressive Regressor, Elastic Net, Lasso Least Angle
Regressor and Least Angle Regression. The above-mentioned
models are implemented using the ‘pycaret’ [37] library
available in python. Ten-fold cross validation is performed to
ensure that the models do not overfit.
k-NN [39] is a powerful algorithm which can be used for
classification and regression cases. In regression, the algorithm
makes use of feature similarity to determine the nearest
neighbours. The nearest neighbours can be found using a
distance formula such as Euclidean or Manhattan distance. k
represents the number of neighbours chosen. The model was
tested for varying K values, the system achieved its lowest
RMSE of 12.17 at 5 neighbours.
The Bayesian Ridge [40] is a probabilistic regression
algorithm. The algorithm makes use of probabilistic
distributors over classical point estimators. The algorithm is
able to tackle datasets that are poorly distributed or have
insufficient data. The algorithm achieved an RMSE of 13.13.
The loss function for the ridge regression algorithm is the linear
least square function, using this function, the Ridge Regression
achieved an RMSE of 13.14.
Linear regression [41] is a statistical model that creates a
linear relationship between the features of the dataset and the
class label. The input features are fed as a linear combination,
along with the class label. The model achieved an RMSE of
13.19.
The Huber regression [42] makes use of the Huber loss
function, which ensures that the model is not heavily affected
by outliers in the dataset, the Huber regression algorithm
achieved an RMSE of 13.4.
TheilSen regression [43] is another statistical algorithm that
computes the best fit line by calculating the median of the
slopes. The algorithm is efficient in tackling random outliers in
the dataset. The algorithm achieved an RMSE of 13.57.
Lasso [44] is an acronym that means Least Absolute
Shrinkage and Selection Operator. As the name suggests, the
lasso regression algorithm makes use of shrinkage. All the data
points are shrunk to a value, which can be the mean or the
central value. Using this, the future values are predicted. The
Fig. 11. Prediction curves for a) Open b) Close c) High and d) Low prices by
Bi-LSTM based Multitask model
algorithm achieved an RMSE of 14.58.
SVM [45] is a supervised algorithm which can be used for
∑ | | classification and regression. The algorithm constructs a
(1) hyperplane and a decision boundary. Any value that lies outside
∑ (2) the decision boundary is an outlier, the hyperplane acts as the
best fit line for the model. The linear kernel is used. The
∑ (3) algorithm achieved an RMSE of 14.90.
If the linear regression model needs to be developed to work
on a higher dimension dataset, the least angle algorithm [46]

0084

Authorized licensed use limited to: ANNA UNIVERSITY. Downloaded on December 26,2020 at 05:22:48 UTC from IEEE Xplore. Restrictions apply.
can be made use. The algorithm achieved an RMSE of 227.5, REFERENCES
since the dataset is not of a higher dimension, the RMSE value [1] B. Chhimwal and V. Bapat, “Impact of foreign and domestic investment
is significantly high. By applying the LASSO algorithm on the in stock market volatility: Empirical evidence from India,” Cogent
least angle algorithm, shrinkage is possible. The lasso least Economics & Finance, vol. 8, no. 1, Apr. 2020.
angle regression achieved an RMSE of 18.6. [2] S. Sridhar, S. Mootha, and S. Subramanian, “Decentralized Stock
Exchange Implementation using Ethereum,” in 2020 International
Seminar on Intelligent Technology and Its Applications (ISITIA), pp. 234-
TABLE III. COMPARISON OF PROPOSED MODELS AGAINST MACHINE 241, 2020.
LEARNING MODELS
[3] C. Pop et al., “Decentralizing the Stock Exchange using Blockchain An
Model MAE MSE RMSE Ethereum-based implementation of the Bucharest Stock Exchange,” in
2018 IEEE 14th International Conference on Intelligent Computer
Bi-LSTM Seq2Seq Model 2.6845 15.899 3.9874 Communication and Processing (ICCP), pp. 459-466, 2018.
Bi-LSTM Multitask [4] N. Sakthivel and A. Saravanakumar, “Investors’ Satisfaction on Online
5.4586 61.9907 7.8734
Learning Share Trading and Technical Problems Faced by the Investors: A Study
k-Neighbours Regressor 9.2576 149.553 12.1733 in Coimbatore District of Tamilnadu,” International Journal of
Management Studies, vol. V, no. 3(9), p. 71, Jul. 2018.
Bayesian Ridge 9.7404 174.392 13.1391
[5] D. Shah, H. Isah, and F. Zulkernine, “Stock Market Analysis: A Review
Ridge Regression 9.7545 174.625 13.1467 and Taxonomy of Prediction Techniques,” International Journal of
Financial Studies, vol. 7, no. 2, p. 26, May 2019.
Linear Regression 9.8032 175.807 13.1914
[6] U. Hathi, “Indian Companies Act 2013 Highlights and Review,” SSRN
Huber Regressor 9.4407 181.94 13.4108 Electronic Journal, 2014.
[7] S. V. Shenoy and K. Srinivasan, “Relationship of IPO Issue Price and
TheilSen Regressor 10.1028 186.304 13.5736 Listing Day Returns with IPO Pricing Parameters,” International Journal
Random Sample Consensus 9.4097 188.366 13.6552 of Management Studies, vol. V, no. 4(1), p. 11, Oct. 2018.
Orthogonal Matching [8] G. Tanty and P. K. Patjoshi, “A Study on Stock Market Volatility Pattern
10.2667 194.284 13.8447 of BSE and NSE in India,” Asian Journal of Management, vol. 7, no. 3,
Pursuit
p. 193, 2016.
Lasso Regression 11.1123 214.692 14.5814
[9] S. Sridhar, S. Mootha, and S. Subramanian, “Detection of Market
Support Vector Machine 10.5572 224.368 14.9036 Manipulation using Ensemble Neural Networks,” in 2020 International
Conference on Intelligent Systems and Computer Vision (ISCV), pp. 1–8,
Passive Aggressive
10.8883 226.655 14.9515 2020
Regressor
[10] S. Zavadzki, M. Kleina, F. Drozda, and M. Marques, “Computational
Elastic Net 12.3847 276.343 16.5724 Intelligence Techniques Used for Stock Market Prediction: A Systematic
Lasso Least Angle Review,” IEEE Latin America Transactions, vol. 18, no. 04, pp. 744–755,
13.8716 348.526 18.6167
Regressor Apr. 2020.
Least Angle Regression 151.758 159.069 227.599 [11] M. C. Joshi, “Factors Affecting Indian Stock Market,” SSRN Electronic
Journal, 2013.
[12] S. Alhazbi, A. B. Said and A. Al-Maadid, "Using Deep Learning to
As can be seem from Table III, the proposed Seq2Seq Predict Stock Movements Direction in Emerging Markets: The Case of
model and Bi-LSTM Multitask model comfortably outperform Qatar Stock Exchange," 2020 IEEE International Conference on
the existing regression models. A low RMSE indicates that the Informatics, IoT, and Enabling Technologies (ICIoT), Doha, Qatar, pp.
440-444, 2020.
model is predicting the future value with accuracy. As can be
[13] D. Wei, “Prediction of Stock Price Based on LSTM Neural Network,” in
seen from Table III, the two proposed models have a 2019 International Conference on Artificial Intelligence and Advanced
significantly low RMSE value in comparison to existing Manufacturing (AIAM), pp. 544-547, 2019.
regression algorithms. This demonstrates that the proposed [14] Y. Keneshloo, T. Shi, N. Ramakrishnan, and C. K. Reddy, “Deep
models are able to predict the upward and downward trend of Reinforcement Learning for Sequence-to-Sequence Models,” IEEE
the price of a stock. Transactions on Neural Networks and Learning Systems, vol. 31, no. 7,
pp. 2469–2489, 2019.
VI. CONCLUSION AND FUTURE WORKS [15] K. Palasundram, N. Mohd Sharef, N. A. Nasharuddin, K. A. Kasmiran,
and A. Azman, “Sequence to Sequence Model Performance for Education
In this paper, a system to predict the future Open, High, Chatbot,” International Journal of Emerging Technologies in Learning
Close, Low (OHCL) price of a stock using a Bi-Directional (iJET), vol. 14, no. 24, p. 56, Dec. 2019.
LSTM based Sequence to Sequence Modelling has been [16] H. Jain and G. Harit, "An Unsupervised Sequence-to-Sequence
Autoencoder Based Human Action Scoring Model," 2019 IEEE Global
proposed. The system is tested on the stock price of Tata Conference on Signal and Information Processing (GlobalSIP), Ottawa,
Consumer Products Limited from the National Stock Exchange ON, Canada, pp. 1-5, 2019.
(NSE) of India. The Seq2Seq model and Multitask model [17] K. A. Althelaya, E.-S. M. El-Alfy, and S. Mohammed, “Evaluation of
achieved an RMSE value of 3.98 and 7.87 respectively. The bidirectional LSTM for short-and long-term stock market prediction,” in
proposed models were compared to various existing regression 2018 9th International Conference on Information and Communication
Systems (ICICS), pp. 151-156, 2018.
models, and outperformed them.
[18] J. Chou and T. Nguyen, "Forward Forecast of Stock Price Using Sliding-
In the future, the system can be tested for a combination of Window Metaheuristic-Optimized Machine-Learning Regression," in
static and dynamic features. Further, sentiment analysis can be IEEE Transactions on Industrial Informatics, vol. 14, no. 7, pp. 3132-
performed on news articles, blogs, social media, etc, to 3142, July 2018.
determine the upward or downward trend of a stock. [19] Y. Zhang and Q. Yang, “An overview of multi-task learning,” National
Science Review, vol. 5, no. 1, pp. 30–43, 2018.

0085

Authorized licensed use limited to: ANNA UNIVERSITY. Downloaded on December 26,2020 at 05:22:48 UTC from IEEE Xplore. Restrictions apply.
[20] B. Ko and H. Choi, "Paraphrase Bidirectional Transformer with Multi- [33] “[Link] — scikit-learn 0.22.1
task Learning," 2020 IEEE International Conference on Big Data and documentation,” [Link], 2019. [Online]. Available:
Smart Computing (BigComp), Busan, Korea (South), pp. 217-220, 2020. [Link]
[21] Y. Bing, J. K. Hao, and S. C. Zhang, “Stock Market Prediction Using [Link]/stable/modules/generated/[Link]
Artificial Neural Networks,” Advanced Engineering Forum, vol. 6–7, pp. .html.
1055–1060, Sep. 2012. [34] L. Lu, X. Zhang, K. Cho, and S. Renals, “A study of the recurrent neural
[22] M. Jia, J. Huang, L. Pang, and Q. Zhao, “Analysis and Research on Stock network encoder-decoder for large vocabulary speech recognition,” in
Price of LSTM and Bidirectional LSTM Neural Network,” in Proceedings of the Annual Conference of the International Speech
Proceedings of the 3rd International Conference on Computer Communication Association, INTERSPEECH, pp. 3249–3253, 2015.
Engineering, Information Science & Application Technology (ICCIA [35] S. V. G. Reddy, K. T. Reddy, and V. ValliKumari, “Optimization of Deep
2019), pp.467-473, 2019. Learning using various Optimizers, Loss functions and Dropout,”
[23] C.-H. Cho, G.-Y. Lee, Y.-L. Tsai, and K.-C. Lan, “Toward Stock Price International Journal of Recent Technology and Engineering (IJRTE),
Prediction using Deep Learning,” in Proceedings of the 12th IEEE/ACM vol. 7, no. 4S2, pp. 448–455, 2018.
International Conference on Utility and Cloud Computing Companion - [36] A. Botchkarev, “A New Typology Design of Performance Metrics to
UCC ’19 Companion, pp. 133–135, 2019. Measure Errors in Machine Learning Regression Algorithms,”
[24] J. Eapen, D. Bein, and A. Verma, “Novel Deep Learning Model with CNN Interdisciplinary Journal of Information, Knowledge, and Management,
and Bi-Directional LSTM for Improved Stock Market Index Prediction,” vol. 14, pp. 045–076, 2019.
in 2019 IEEE 9th Annual Computing and Communication Workshop and [37] “Home,” PyCaret. [Online]. Available: [Link]
Conference (CCWC), pp. 0264-0270, 2019. [38] S. Ravikumar and P. Saraf, "Prediction of Stock Prices using Machine
[25] C. Li, D. Song, and D. Tao, “Multi-task Recurrent Neural Networks and Learning (Regression, Classification) Algorithms," 2020 International
Higher-order Markov Random Fields for Stock Price Movement Conference for Emerging Technology (INCET), Belgaum, India, pp. 1-5,
Prediction,” in Proceedings of the 25th ACM SIGKDD International 2020.
Conference on Knowledge Discovery & Data Mining, pp. 1141–1151, [39] J. Tanuwijaya and S. Hansun, “LQ45 Stock Index Prediction using k-
2019. Nearest Neighbors Regression,” International Journal of Recent
[26] W. Khan, M. A. Ghazanfar, M. A. Azam, A. Karami, K. H. Alyoubi, and Technology and Engineering, vol. 8, no. 3, pp. 2388–2391, Sep. 2019.
A. S. Alfakeeh, “Stock market prediction using machine learning [40] Y. Yang and Y. Yang, "Hybrid Prediction Method for Wind Speed
classifiers and social media, news,” Journal of Ambient Intelligence and Combining Ensemble Empirical Mode Decomposition and Bayesian
Humanized Computing, Mar. 2020. Ridge Regression," in IEEE Access, vol. 8, pp. 71206-71218, 2020.
[27] “TATA CONSUMER PRODUCTS LIMITED Share Price Today, Live [41] F. Mar'i, U. Pratiwi, I. Oktanisa and F. Utaminingrum, "Comparative
NSE Stock Price, News – NSE India,” NSE India, 2020. [Online]. Study of Numerical Methods in Multiple Linear Regression For Stock
Available: [Link] Prediction Jakarta Islamic Index (JII)," 2019 International Conference on
quotes/equity?symbol=TATACONSUM. [Accessed: 17-Sep-2020]. Sustainable Information Engineering and Technology (SIET), Lombok,
[28] S. Hwang, G. Jeon, J. Jeong, and J. Lee, “A Novel Time Series based Indonesia, pp. 110-115, 2019.
Seq2Seq Model for Temperature Prediction in Firing Furnace Process,” [42] Q. Sun, W.-X. Zhou, and J. Fan, “Adaptive Huber Regression,” Journal
Procedia Computer Science, vol. 155, pp. 19–26, 2019. of the American Statistical Association, vol. 115, no. 529, pp. 254–265,
[29] A. Joshi, K. Mehta, N. Gupta and V. K. Valloli, "Data Generation Using Apr. 2019.
Sequence-to-Sequence," 2018 IEEE Recent Advances in Intelligent [43] A. Yenkikar, M. Bali, and N. Balu, “EMP-SA: Ensemble Model based
Computational Systems (RAICS), Thiruvananthapuram, India, pp. 108- Market Prediction using Sentiment Analysis,” International Journal of
112, 2018. Recent Technology and Engineering, vol. 8, no. 2, pp. 6445–6452, Jul.
[30] G. Aalipour, P. Kumar, S. Aditham, T. Nguyen and A. Sood, 2019.
"Applications of Sequence to Sequence Models for Technical Support [44] S. S. Roy, D. Mittal, A. Basu, and A. Abraham, “Stock Market
Automation," 2018 IEEE International Conference on Big Data (Big Forecasting Using LASSO Linear Regression Model,” in Advances in
Data), Seattle, WA, USA, pp. 4861-4869, 2018. Intelligent Systems and Computing, pp. 371–381, 2015.
[31] Y. Chen, W. Lin and J. Z. Wang, "A Dual-Attention-Based Stock Price [45] B. M. Henrique, V. A. Sobreiro, and H. Kimura, “Stock price prediction
Trend Prediction Model With Dual Features," in IEEE Access, vol. 7, pp. using support vector regression on daily and up to the minute prices,” The
148047-148058, 2019. Journal of Finance and Data Science, vol. 4, no. 3, pp. 183–201, Sep.
[32] G. Balikas, S. Moura, and M.-R. Amini, “Multitask Learning for Fine- 2018.
Grained Twitter Sentiment Analysis,” in Proceedings of the 40th [46] C. Fraley and T. Hesterberg, “Least Angle Regression and LASSO for
International ACM SIGIR Conference on Research and Development in Large Datasets,” Statistical Analysis and Data Mining, vol. 1, no. 4, pp.
Information Retrieval, pp. 1005–1008, 2017. 251–259, Mar. 2009.

0086

Authorized licensed use limited to: ANNA UNIVERSITY. Downloaded on December 26,2020 at 05:22:48 UTC from IEEE Xplore. Restrictions apply.

Common questions

Powered by AI

The window length impacts the Seq2Seq model's effectiveness by determining the size of the data sequence used for predictions. A well-chosen window length minimizes the RMSE of the prediction model, as shown in the Bi-LSTM based Seq2Seq model with window size optimization at 60 days, providing an improved balance between historical data use and computational efficiency .

Multitask learning improves prediction accuracy by modeling the interrelation of OHLC prices as subtasks. It enables simultaneous prediction of these prices, allowing knowledge from each subtask to be shared with a common model that predicts the output sequence. This joint learning process helps reduce prediction error by leveraging overlapping information among OHLC prices .

The sliding window approach is significant as it dynamically creates sequences of OHLC prices used as inputs for prediction models, which helps capture temporal dependencies and reduce prediction errors. By optimizing window size, the approach minimizes the RMSE, as shown in both Seq2Seq and multitask Bi-LSTM models, leading to more accurate predictions .

The dataset preprocessing involved cleaning the data to remove null values and normalizing it using MinMaxScaler. This scaling helped adjust data features into a fixed range, making it suitable for processing by machine learning models. The dataset, consisting of OHLC prices, was then split into training and testing sets at a 75:25 ratio .

Seq2Seq modeling helps forecast the consecutive output of a given input sequence by mapping sequences of variable length to an output sequence. The underlying architecture involves an encoder-decoder framework where the encoder maps the input sequence to a fixed-length vector, which is then decoded to predict the output sequence. This architecture is implemented using Bi-directional LSTM units to handle OHLC price sequences from stock data .

AI and machine learning play a pivotal role by enabling sophisticated modeling techniques capable of processing complex and high-dimensional data from various sources, such as OHLC sequences, news, and social media. They allow for enhanced prediction through model architectures like Seq2Seq and multitask learning, which simultaneously manage multiple variables and leverage vast computational power to deliver precise forecasts .

Stock prices are influenced by numerous external factors such as weather forecasts, political speeches, social media, news, and global economic trends. AI can address these complexities by processing vast amounts of unstructured data from multiple sources using advanced algorithms and high-performance computing, allowing investors to make informed predictions despite these uncontrollable variables .

Bi-directional LSTM offers the advantage of processing data sequences in both forward and reverse directions, which enables the model to capture long-term dependencies and improve the handling of sequence length variability. This enhances the model's ability to predict stock prices by understanding complex dependencies in the data sequence .

Recurrent Neural Networks provide the benefit of capturing temporal dependencies in time-series stock data, crucial for pattern recognition in sequential data sets like OHLC prices. Challenges include their vulnerability to vanishing gradients, which can hinder learning sequences with long-term dependencies. Advanced forms, like LSTM and Bi-LSTM, mitigate these issues but add computational complexity .

While not directly covered in the provided sources, dual attention mechanisms generally enhance prediction accuracy by allowing models to focus on key features from both the input and output sequence spaces. In the context of stock prediction, such mechanisms enable better capture of critical signals from OHLC data, improving the interpretability and precision of the forecasted results.

You might also like