0% found this document useful (0 votes)
6 views16 pages

Deep Learning

This research article investigates hybrid deep learning models for sentiment analysis, combining LSTM, CNN, and SVM to enhance accuracy across various datasets. The study demonstrates that these hybrid models outperform single models in sentiment classification, particularly in diverse domains such as social media and reviews. The findings suggest that hybrid approaches can effectively address challenges in sentiment analysis, leading to improved reliability and efficiency.

Uploaded by

Wiem Jouini
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views16 pages

Deep Learning

This research article investigates hybrid deep learning models for sentiment analysis, combining LSTM, CNN, and SVM to enhance accuracy across various datasets. The study demonstrates that these hybrid models outperform single models in sentiment classification, particularly in diverse domains such as social media and reviews. The findings suggest that hybrid approaches can effectively address challenges in sentiment analysis, leading to improved reliability and efficiency.

Uploaded by

Wiem Jouini
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Hindawi

Complexity
Volume 2021, Article ID 9986920, 16 pages
[Link]

Research Article
Hybrid Deep Learning Models for Sentiment Analysis

Cach N. Dang ,1,2,3 Marı́a N. Moreno-Garcı́a,2 and Fernando De la Prieta3


1
Department of Information Technology, Ho Chi Minh City University of Transport (UT-HCMC), Ho Chi Minh 70000, Vietnam
2
Data Mining (MIDA) Research Group, University of Salamanca, Salamanca 37007, Spain
3
Biotechnology, Intelligent Systems and Educational Technology (BISITE) Research Group, University of Salamanca,
Salamanca 37007, Spain

Correspondence should be addressed to Cach N. Dang; cach@[Link]

Received 2 April 2021; Accepted 6 August 2021; Published 13 August 2021

Academic Editor: Tao Jia

Copyright © 2021 Cach N. Dang et al. This is an open access article distributed under the Creative Commons Attribution License,
which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Sentiment analysis on public opinion expressed in social networks, such as Twitter or Facebook, has been developed into a wide
range of applications, but there are still many challenges to be addressed. Hybrid techniques have shown to be potential models for
reducing sentiment errors on increasingly complex training data. This paper aims to test the reliability of several hybrid techniques
on various datasets of different domains. Our research questions are aimed at determining whether it is possible to produce hybrid
models that outperform single models with different domains and types of datasets. Hybrid deep sentiment analysis learning
models that combine long short-term memory (LSTM) networks, convolutional neural networks (CNN), and support vector
machines (SVM) are built and tested on eight textual tweets and review datasets of different domains. The hybrid models are
compared against three single models, SVM, LSTM, and CNN. Both reliability and computation time were considered in the
evaluation of each technique. The hybrid models increased the accuracy for sentiment analysis compared with single models on all
types of datasets, especially the combination of deep learning models with SVM. The reliability of the latter was
significantly higher.

1. Introduction the shortcoming of short text in deep learning models.


Besides, the study by Qian et al. [16] revealed that LSTM
Sentiment analysis on information from social networks, behaves efficiently when used on different text levels of
such as Twitter or Facebook, is a research topic of growing weather-and-mood tweets. After reviewing some recent
interest today. Although much work has been done in this studies [1, 11, 12, 15, 17–20], we found that CNN and
area, there are still many challenges to be addressed, in- RNN are outperforming methods with a relatively high
cluding improving model reliability, reducing processing overall accuracy. Both shallow neural networks and deep
time, and applying techniques developed for specific types of neural networks are capable of approximating any
data and specific data domains [1]. In recent years, deep function. However, when contrasted to shallow neural
learning models have been extensively applied in the field of networks, deep neural networks have the advantage of + deep
sentiment analysis, where their great potential has been being able to do the feature extraction in the process of neural
network
demonstrated. learning on large datasets. This is primarily because the
Several studies are focused exclusively on building a deep models are able to extract/build better features than
single model from a single (or some) dataset(s) in a shallow models, using the intermediate hidden layers to
particular domain, such as marketing strategies [2], fi- achieve this [21, 22]. For the same level of accuracy, deep
nancial forecasting [3–5], and medical analysis [6, 7]. For neural networks can be much more efficient in terms of
social network applications, sentiment polarity-based computation and number of parameters. Deep neural
deep learning applied to tweets is described thoroughly in networks are able to create deep representations; at every
[8–14]. Hassan and Mahmood [15] proved that CNN and layer, the network learns a new, more abstract repre-
recurrent neural networks (RNN) models can overcome sentation of the input.
2 Complexity

Although a single machine learning method is relatively regardless of the types of social network datasets; providing
reliable when applied within certain domains, each deep an experimental study to evaluate the performance of hybrid
learning approach has its own advantages and disadvan- deep learning models; and detailing a performance com-
tages. LSTM normally yields better results but requires more parison of sentiment analysis methods with some state-of-
processing time than CNN, and CNN requires fewer art methods.
hyperparameters and less supervision. Meanwhile, the The paper is organized as follows. Section 2 presents an
LSTM performs more accurately for long sentences but overview of related work; Section 3 describes the method-
requires a longer time to process [1]. ology in this research area; Section 4 contains the proposed
The approach of combining two (or more) methods is hybrid models; Section 5 describes and discusses the results
introduced [23–25] as a means of incorporating the ad- of our experiments; and Section 6 offers our conclusions.
vantages of both and thus fills some shortcomings of in-
dividual methods. Alfrjani et al. [25] combined machine 2. Related Work
learning and semantic knowledge base for improving ac-
curacy of sentiment analysis on reviews (improvement 1% to The purpose of this study is to build hybrid models for
6%). In another case, Gupta and Joshi [23] proposed a hybrid sentiment analysis that can improve accuracy. We have
method that combines lexicon and machine learning for previously examined the methods proposed and applied in
sentiment analysis on tweets (improvement 2% to 6%). A other studies, which are discussed as follows.
hybrid system with collaborative functions, therefore, is There are many ways to build hybrid models. In [26–28],
better able to address potential pitfalls, if any exist, asso- the authors combined a CNN model and SVM that can
ciated with one single system. The effectiveness of the in- improve the accuracy in image recognition. A convolutional
tegrated models may vary based on different tasks. The CNN network layer is used for extracting features and SVM
enhanced by SVM [26–28], CNN with RNN [29–32], and functions as a recognizer. Original CNN is used with
Lexicon-based analysis with machine learning [33, 34] Softmax functions. Srinidhi et al. [36] proposed a hybrid
showed an enhanced result. The combination of CNN, model that combined LSTM and SVM with a radial basis
LSTM, and SVM aims to take advantage of the two deep function kernel for the textual classification of positive and
network architecture models and SVM algorithms when negative sentiments. The hybrid model was evaluated on the
performing sentiment analysis on different domains and IMDb movie review datasets. These models are combined
types of datasets. Moreover, there are different types of input from single deep learning models with SVM for classifica-
data obtained from social networks, such as tweets and tion. Some of them are applied for image recognition. Our
reviews. Within and across these types, the input data also research combines two deep learning models and then uses
contains differences, for example, the distribution of the SVM or ReLU for classification.
lengths of the tweets and reviews, the diversity of topics in Akhtar et al. [37] built a hybrid deep learning archi-
each dataset, the sample size, and the greater or lesser tecture, which is highly efficient for sentiment analysis in
presence of explicit sentiments and irrelevant information. resource-poor languages. They used CNN for learning
Some approaches may be unable to perform well in different sentiment embedded vectors and SVM for sentiment clas-
domains, with inadequate accuracy and performance in sification. The model was tested on four Hindi datasets
sentiment analysis [1, 35]. As a result, certain approaches covering varied domains. Vo et al. [31] used a multichannel
may be ill-suited and difficult to apply to certain types of LSTM-CNN model for sentiment analysis on reviews/
input data. comments from e-commerce sites. In addition, hybrid
A question raised in our study is whether hybrid models CNN–LSTM models are applied for sentiment analysis on
perform better than single models regardless of the char- movie reviews by Rehman et al. [30]. The same techniques
acteristics of the datasets. Therefore, our work examines how are used in several works, for example, [29, 38–40]. Kaur
selected hybrid models behave with different types of et al. [41] designed an algorithm called a hybrid heteroge-
datasets from different domains. In this work, we evaluated neous support vector machine (H-SVM). They performed
and validated the combination of three models CNN, LSTM, sentiment analysis on Twitter data related to COVID-19.
and SVM. We considered the relationship between models Kastrati et al. [42] employed three different deep learning
and its advanced capacities to extract characteristics, store models such as CNN, LSTM, and CNN-LSTM for classifying
past information and nodes, and classify text. First, in the Facebook comments related to the COVID-19 pandemic.
initial stages of the model, two possible variations in the They used pretrained word embedding method called
sequence of CNN and LSTM are introduced. Then, for each FastText (an extension to Word2vec proposed by Facebook
of these alternatives, two new variations are introduced: the in 2016) and a contextualized word embedding model,
use of CNN with ReLU function or SVM. We applied these BERT, to learn and generate word vector. Both research
models with word embedding on eight datasets, including scored tweet/comment as positive, negative, or neutral.
tweets and reviews. The results of our experiments showed However, these models were individually tested on different
that the combined models increased the accuracy of senti- datasets in a particular domain or tested on few sample
ment analysis. datasets. Therefore, their validity is not generally proven.
This paper offers three important contributions to the A study by Jnoub et al. [19] focused on providing a
literature by highlighting four hybrid deep learning models generalized model for sentiment analysis that combined
for sentiment analysis that results in improved accuracy CNN with their own algorithm to transform reviews to
Complexity 3

vectors. The model was evaluated on three different datasets: were considered for the selection including the ability to
IMDb, movie reviews, and their own dataset collected from avoid privacy concerns [52], acceptance in the research
Amazon reviews. Ombabi et al. [43] proposed a hybrid deep community, diversity of sources and topics, and size. The
learning model that combines CNN and LTSM. In addition, selected datasets enable a comprehensive comparison of the
FastText is used for word embedding and SVM for classi- sentiment analysis approaches examined in this paper. The
fication in the Arabic language. In our work, both Word2vec aim of the experiment is to understand whether the models
and BERT were applied for word embedding. We proposed give consistently accurate results regardless of the dataset
four types of hybrid deep learning models based on CNN, type and size.
LSTM, and SVM for classifying both tweets and reviews. The experiments were conducted using eight datasets.
Furthermore, other studies combine Lexicon-based Three datasets contain tweets (Sentiment140, Tweets Airline,
analysis with machine learning [33, 34] or sentiment lexi- and Tweets SemEval) and five datasets contain reviews
cons and polarity shifting devices [44]. The research by (IMDb movie reviews (1) and (2) and Cornell movie review).
Sánchez-Rada and Iglesias [24] deals with the problem of Among the tweets datasets, Sentiment140 [53], the largest,
user and content sentiment classification. They proposed a has 1.6 million tweets, each one labelled as either positive or
hybrid model that merges features from different levels of negative sentiment, while the others, Tweets Airline [54] and
social context. The model is evaluated in different datasets. A Tweets SemEval [55], contain 14,640 and 17,750 tweets,
study from Wang et al. [45] presented a hybrid approach, in respectively, labelled as positive, negative, or neutral. The
which sentiment analysis of reviews about movies is used to five review datasets include a total of 125,000 comments
improve a preliminary recommendation list obtained from from user reviews of movies (IMDb movie reviews (1) [56],
the combination of collaborative filtering and content-based IMDb movie reviews (2) [57], and Cornell movie reviews
methods. In the same approach, the use of a sentiment [58]), books, and music (book and music reviews [59]),
classifier induced from movie reviews as a second filter after labelled as either positive or negative sentiments. They are
collaborative filtering was proposed by Pandey et al. [10]. discussed in more detail in [1].
These research projects use traditional techniques to per- After examining the collected datasets, we saw that six
form sentiment analysis. Our research applies deep learning out of eight datasets are initially labelled as positive and
techniques for improving the accuracy of sentiment negative, and the sample on each label is relatively equal. The
classification. two datasets Airline and Tweet SemEval contain not only
Recently, transfer learning has been successfully applied positive and negative labels but also neutral label. Having a
in sentiment analysis, in which lower network layers are balanced class distribution is important to ensure that prior
trained on high-resource supervised datasets, such as BERT probabilities are not biased for training models and doing
(proposed by researchers at Google AI language in 2018 classification [60]. In this research, we focus on polarity
[46]) and XLNET [47]. Examples can be found in [48–51], sentiment analysis, based on two classes positive and neg-
where BERT and XLNET were applied for sentiment anal- ative. The size of these datasets was reduced by removing the
ysis. The evaluation of different datasets and languages neutral labels. The remaining positive and negative classes
provides significant results. However, it also requires suf- are readjusted to be balanced. In addition, we applied k-fold
ficiently powerful hardware, large datasets, and long pro- cross-validation to the data in order to evaluate the models.
cessing times when applying these techniques. For example, In this way, the tests cover all instances of the datasets
BERT-Base model has 110M parameters, and BERT-Large avoiding bias towards a particular subset of the data. Table 1
model has 340M parameters: pretraining is fairly expensive, shows the number of samples (positive and negative) taken
requiring four days on 4 to 16 cloud TPUs. from each of dataset for performing experiments.

3. Methodology
3.2. Preprocessing and Building the Feature Vector.
Considering all of the advantages and potential of hybrid Sentiment classification can be carried out on three levels of
models and aiming at improving the performance of sen- extraction: the document, sentence, and aspect or feature
timent analysis techniques, our paper evaluates four hybrid [61]. In our experiments, we applied document-based
models. The methodology is focused on three main com- sentiment analysis with word embedding techniques on
ponents: the data to be used; process to build the feature eight datasets of tweets and reviews. Sentiment analysis
vectors; building of hybrid methods for an appropriate requires that text-training data be cleaned before using as
sentiment analysis solution. These algorithms are applied to input for classification models. Irrelevant information in text
predict the sentiment polarity of the text and classify it or sentence data, including white space, punctuation, and
according to that polarity. stop words, is removed. Two techniques commonly used for
this task are TF-IDF and word embedding. Our proposal
uses the latter because it provides better results than TF-IDF
3.1. Datasets. Our study does not focus on solving a problem [1]. We then used word embedding models, BERT and
in a particular domain but on providing an evaluation for Word2vec, to build the feature vector.
general application models. In this study, we used several BERT is a language model for nature language pro-
public datasets instead of generating and labelling new cessing, and it was published by researchers at Google AI
datasets of a specific application domain. Multiple criteria Language in 2018 [46]. BERT was developed after Word2vec
4 Complexity

Table 1: Number samples of datasets. the dataset. This is commonly done in other works
# Datasets Number of samples
[30, 64–66].
The fixed length d is selected as follows: datasets related
1 Sentiment140 (10%) 160.000
to tweets usually have a small length variation due to the
2 Tweets Airline 4.726
3 Tweets SemEval 9.300 limit of tweets to a maximum of 280 characters; thus, this
4 IMDb movie reviews (1) 50.000 fixed length d is chosen to be the maximum length of the
5 IMDb movie reviews (2) 25.000 sample in the dataset. For the remaining datasets, the length
6 Cornell movie reviews 10.662 d is selected from 300 to 500, based on the histogram of every
7 Book reviews 2.000 single dataset. It could be possible to take a fixed length d,
8 Music reviews 2.000 instead of different lengths, for tweets and reviews. However,
if set length d is larger, it will waste much memory, and if set
length d is smaller, it will miss some review data.
and includes some advances over Word2vec, such as support
for out-of-vocabulary (OOV) words.
Word2vec was published in 2013 by Tomas Mikolov at 3.3. Hybrid Methods. There are numerous methods to build
Google [62]. This unsupervised learning model has trained up a hybrid model for sentiment analysis. In this study, we
datasets from a large corpus. The dimension of Word2vec is tested the combination of several successful approaches. As
much less than the dimension of one-hot encoding, with a shown in Figure 3, we start by using Word2vec or a pre-
matrix NxD, with N being the number of documents and D trained BERT model to create the feature vector. We then
being the dimension of word embedding. Word2vec con- vary the order of the CNN and LSTM models used in the
tains two models: skip-gram and continuous bag-of-words next stages: Word2vec/BERT - > CNN - > LSTM or
(CBOW). Both models are based on the probability of words Word2vec/BERT - > LSTM - > CNN. We also vary the final
occurring in proximity to each other. Skip-gram allows us to stage of the model, using a ReLU function or using an SVM.
start with a word and predict words that likely surround it. Combining these two types of variation yields the four
However, one of the major drawbacks of using Word2vec is hybrid approaches that we have tested:
a lack of support for out-of-vocabulary words. To work
around this issue, we use the special token [UNK] for words (1) Word2vec/BERT - > CNN - > LSTM - > Relu
not found in the vocabulary. In addition, we also retrain the (2) Word2vec/BERT - > LSTM - > CNN - > Relu
Word2vec model according to our vocabulary datasets with (3) Word2vec/BERT - > CNN - > LSTM - > SVM
all words that appear more than five times, reducing the use
(4) Word2vec/BERT - > LSTM - > CNN - > SVM
of the special token.
One issue in conducting sentiment analysis modelling is Two approaches were used in our experiments to create
the varying length of the samples of the dataset. While deep feature vectors. The first approach was Word2vec initialized
learning models require fixed input vectors. Figures 1 and 2 with random weights to learn the embedding for all words in
show histograms of the datasets of reviews and tweets after our training datasets. Because Word2vec does not include
they were cleaned. The x-axis represents the length of the contextual analysis to handle complex semantical or poly-
data samples, and the y-axis is the frequency of appearance. morphic cases in natural languages, our second approach
Some histograms are rather ragged because we chose dif- was BERT. A pretrained BERT model was used in this study.
ferent types of datasets from different sources. Standardizing After adjusting the parameters, the BERT model was used as
data by smoothing outlines based on sample size could well a feature extractor to generate input data for the proposal of
fit the models [63]. In this study, we keep nearly raw data for hybrid models. The tweets and reviews data were fed into the
sentiment analysis with the purpose of creating the necessary BERT model to generate the feature vectors, which are the
conditions to compare the efficiency of other models. input to the hybrid models that perform the classification.
We can see in Figures 1 and 2 that the data samples are The next step combines CNN and LSTM deep learning
quite widely varied in length. Therefore, it is necessary to set models, which are used because of their good performance
the data samples to the same length. The conversion of data on sentiment analysis [1], as well as taking advantage of the
samples adjusted to the same length is done as follows. two network architectures when performing sentiment
For each dataset, we select a fixed length called d; for analysis on data in different domains. A CNN is a type of
samples shorter than d, we add zeros to the end of the vector. feedforward neural network, since it is composed of multiple
And vice versa, in samples with length greater than d, the layers that process and pass information in one direction,
back will be cut off. However, truncating the length of the from input to output, without cycles. It has a deep neural
data sample will result in a loss of information used in the network architecture [67], typically starting with convolu-
classification process, so it is pivotal to choose a fixed length tional and pooling/subsampling layers that transform inputs
d to minimize truncation of the data samples. In this study, that feed into a fully connected classification layer. In this
we used both tweets and reviews datasets for our proposed research, a single convolutional (1D CNN) was used. LSTM
models. We truncated any tweet or review if its length is is one of the many variations of the RNN architecture [68].
longer than the length of the feature vector. The length of the The LSTM block consists of three so-called gates, the forget
feature vector is chosen to be close to the maximum length of gate, input gate, and output gate, in addition to the input and
tweets and reviews, so very few samples were truncated in output blocks and the memory cell. CNNs are good at
Complexity 5

4000 700
Cornell Movie Reviews
3500 600
3000
500
IMDB Movie reviews
2500
400
2000
300
1500
200
1000

500 100

0 0
0 200 400 600 800 1000 0 10 20 30 40 50

600 250

500 Book reviews Music reviews


200

400
150
300
100
200

50
100

0 0
0 200 400 600 800 1000 1200 0 200 400 600 800 1000
Figure 1: Histograms for different length data samples of reviews datasets.

18000
1200 Tweets Semeval 16000 Sentiment 140 (10%)
14000
1000
12000
800
10000
600 8000
6000
400
4000
200
2000
0 0
0 5 10 15 20 25 30 35 0 5 10 15 20 25 30 35

350
Tweets Airline
300

250

200

150

100

50

0
0 5 10 15 20 25 30 35
Figure 2: Histograms for different length data samples of tweets datasets.
6 Complexity

.1
n1 ReLU
tio
r ia
Va

Va
r ia

n1
tio
n1 SVM

tio
CNN LSTM Fully .2

r ia
connected Output

Va
layer
Bert/
Word2vec

Dataset

Va
.1 ReLU
n2

r ia
tio
r ia

tio
Va

n2
Feature
Var
vector iat ion SVM
2.2
LSTM CNN Fully
connected
layer
Figure 3: Process of methodology for sentiment analysis.

dealing with spatially related data while the RNNs are good
Table 2: A hybrid CNN-LSTM model.
at temporal signals. LSTM can remember forward infor-
mation of the sequence, and multilayer CNN can catch and Layer (type) Output shape Param #
learn local information sufficiently. So, the combination embedding_1 (embedding) (None, 38, 100) 1,808,900
makes use of the best of both worlds, the spatial and tem- conv1d_1 (Conv1D) (None, 38, 512) 154,112
poral worlds. conv1d_2 (Conv1D) (None, 38, 256) 393,472
The final stage is classification. We use the activate conv1d_3 (Conv1D) (None, 38, 128) 98,432
function of ReLU instead of Sigmoid because of the high lstm_1 (LSTM) (None, 500) 1,258,000
convergence. In addition, SVM was chosen for classification dense_1 (dense) (None, 128) 64,128
dense_2 (dense) (None, 128) 16,512
because of its efficiency in word processing, especially in
dense_3 (dense) (None, 1) 129
high dimensional contexts, such as natural language pro- Total params: 3,793,685
cessing. Support vector machine [69] is a supervised ma- Trainable params: 1,984,785
chine learning algorithm that can be used for both Nontrainable params: 1,808,900
classification and regression tasks. It has been widely
exploited with positive results in many areas. In our re-
search, we have applied linear SVMs for classification with Table 3: A hybrid LSTM-CNN model.
the proposed hybrid deep learning models. We extracted Layer (type) Output shape Param #
feature vectors from the top hidden layer and fed it to SVM
embedding_2 (embedding) (None, 38, 100) 1,808,900
that will classify for prediction (“positive” and “negative”). lstm_2 (LSTM) (None, 38, 500) 1,202,000
conv1d_3 (Conv1D) (None, 38, 512) 768,512
4. Proposed Hybrid Models conv1d_5 (Conv1D) (None, 38, 256) 393,472
conv1d_6 (Conv1D) (None, 38, 128) 98,432
In this section, we proposed four hybrid deep learning models flatten_1 (flatten) (None, 4864) 0
on variations in the use of CNN and LSTM in deep learning dense_4 (dense) (None, 128) 622,720
layers and variations of CNN and SVM in the classifier layers. dense_5 (dense) (None, 128) 16,512
The architecture of these hybrid models is shown in Tables 2 dense_6 (dense) (None, 1) 129
and 3, and the details are discussed as follows. Total params: 4,910,677
Trainable params: 3,101,777
Nontrainable params: 1,808,900
4.1. Scenario Combination 1. The first hybrid model com-
bines CNN and LSTM models. The visualization of the feeding it into next deep learning layer. The second layer of
model connection, the connection process, and the data the hybrid model is the LSTM, which produces a 1 × 500
processing flow are indicated in Table 2. matrix that is fed into the classifier. Next, the hybrid model’s
The function embedding is the embedding layer that is classifier is composed of two continuous, fully connected
initialized with random weights, which will learn the em- layers with 128 nodes and, finally, the output layer with a
bedding for all words in the training datasets. The first layer ReLU activation function.
of the hybrid model is the CNN, which receives the vector
produced by word embedding. It has three convolution
layers consisting of 512, 256, and 128 filters, respectively, 4.2. Scenario Combination 2. The second hybrid model
with a kernel size � 3, which receive and process data before combines LSTM and CNN models. The visualization of the
Complexity 7

model connection, the connection process, and the data k models in the cross-validation are induced from training
processing flow are indicated in Table 3. sets of the same size and that the k test sets in all validations
The input data is preprocessed to reshape data for the are also of the same size. It is recommended to split data into
embedding matrix. The first layer of the hybrid model is the equal samples, so that the performance of the models is
LSTM layer. That output has a matrix 13 × 500 and is fed into equivalent.
the second model of the hybrid deep learning model. The
next layer of the hybrid model is the CNN. It has three
5.2. Results. The results of eight sets of experiments are
convolution layers consisting of 512, 256, and 128 filters,
shown: three baseline models (SVM, CNN, and LSTM) and
respectively, with a kernel size � 3, which are in charge of
four hybrid models: CNN and LSTM, LSTM and CNN,
receiving and processing data before feeding it into the next
CNN-LSTM and SVM, LSTM-CNN and SVM referred to as
layer. The CNN output is flattened and transferred to a fully
C-LSTM (or C-L), L-CNN (or L-C), CLSTM-SVM (or
connected layer. Finally, the hybrid model’s classifier is a
CL-S), LCNN-SVM (or LC-S), respectively. A comparison
CNN composed of two continuous fully connected layers
analysis between the results obtained from the proposed
with 128 nodes and the ReLU activation function as the
hybrid methods against the baseline methods is also
output layer.
included.
Our experiments were run twice: once using Word2vec
4.3. Scenario Combinations 3 and 4. Our final hybrid model to train word embedding and once using a pretrained BERT
is based on the hybrid models from scenarios 1 and 2. We model to train word embedding. The results were consis-
used the deep learning stages from those models (CNN- tently better when BERT was used, so Tables 4–8 provide
LSTM and LSTM-CNN) but replaced the classifier. While details on the experimental results using Word2vec and
there are multiple alternatives to the CNN-based ReLU BERT. Figures 4–8 illustrate the comparative results ob-
function used, we have chosen to use SVM for the re- tained with Word2vec and BERT using side-by-side bar
placement classifier. Scenario 3 is based on CNN-LSTM, and charts.
Scenario 4 is based on LSTM-CNN. An architectural The accuracy results shown in Table 4 are very high for
overview of the model is shown in Tables 2 and 3. all datasets and classification models when using a pre-
trained BERT model to extract a feature vector, around 90%,
5. Experimental Results especially, 92.9% in Tweets Airline, and 93.4% in IMDb
movie reviews (1). Moreover, the results prove that hybrid
In this section, we present the experiments conducted to models show higher (or equal) accuracy than single deep
compare the performance of the proposed hybrid models. learning models (SVM, CNN, or LSTM) for seven out of
Moreover, we also examine other common deep learning eight datasets. Regarding the use of Word2vec in the music
models (SVM, CNN, and LSTM). All of them were tested review and book review datasets, CNN’s accuracy results
with the eight datasets introduced in subsection 3.1 that have given in Table 4 are 76.4% and 76.5%, respectively. By
been preprocessed with text processing techniques. Accu- comparison, when using the LCNN-SVM model, the results
racy, AUC, and F-score were the metrics used to evaluate the significantly improve to 83.7% and 82.7%, which represent
performance of the models through all experiments. Since an improvement of 7.3% and 6.2%, respectively.
F-score is derived from recall and precision, we also show For the F-score (Table 7), hybrid models provided higher
these two measures for reference purposes. The results are (or equal) values than single deep learning models for seven
shown, discussed, and analysed in Sections 5.2 and 5.3. out of eight datasets. Regarding the AUC value in Table 8,
the hybrid models also perform better than the single deep-
learning models. The hybrid models using SVM for classi-
5.1. Performance Comparison. Before performing the ex-
fication achieved the best results for six out of eight datasets
periments, the configuration of related parameters, hard-
using Word2vec. Among the datasets, the Tweets Airline
ware devices, and the necessary library facilities were carried
dataset and IMDb movie reviews (1) are the datasets that
out. We used Google Colab Pro with GPU Tesla P100-PCIE-
show the highest values for all metrics in all cases. Book
16GB or GPU Tesla V100-SXM2-16GB [70] and the Keras
reviews and music reviews work well with hybrid LSTM-
[71] and TensorFlow libraries [72]. In all the experiments, we
CNN and LCNN-SVM models. The Sentiment140 dataset
configured the parameter for our code, such as echoes � 4,
has low accuracy in all models. In Figures 1 and 2, we can see
k-fold � 10, and batch size � 32 with reviews and 128 with
the distribution of the total number of samples with the
tweets. The common values for K-fold validation method are
length data sample in the dataset. The Sentiment140 dataset
k � 3, k � 5, and k � 10, and by far, the most popular value
is also different from the other datasets. Number samples of
used in applied machine learning to evaluate models is
length data are so much different.
k � 10. The latter value is used when the dataset is large
enough for the subsets to have a significant number of
examples. This is the case of the datasets used in this work. 5.3. Discussion. As seen in Figures 4 to 8, using pretrained
Thus, nine parts are used as training set and one as test set in BERT produces better results than using Word2vec for
each of the 10 validations. The value of k is chosen to ensure sentiment analysis with all models and datasets. Focusing on
that each train or test sample is large enough to represent the the results of hybrid models, we see that, for each dataset, the
dataset. Furthermore, this procedure ensures that the best results are given by a hybrid model. Hybrid models
8 Complexity

Table 4: Accuracy comparison for different types of datasets.


Word2vec (%) BERT (%)
Datasets
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S
Sentiment140 (10%) 74.2 79.7 80.3 80.1 79.9 80.6 79.9 82.4 84.0 84.1 84.0 84.1 83.5 83.9
Tweets Airline 82.2 86.8 87.1 88.0 87.5 88.6 87.8 92.0 92.9 92.8 92.8 92.8 92.9 92.7
Tweets SemEval 80.5 84.4 86.4 85.0 86.2 85.6 85.7 91.2 91.9 91.8 91.9 91.8 91.7 91.8
IMDb movie reviews (1) 78.9 87.6 88.5 89.7 90.0 90.0 90.3 93.3 93.4 93.4 93.4 93.4 93.4 93.4
IMDb movie reviews (2) 82.8 87.3 85.1 89.2 89.4 89.4 89.4 90.5 90.7 90.6 90.7 90.7 90.7 90.7
Cornell movie reviews 67.7 72.4 76.1 73.0 76.4 76.2 75.9 85.3 87.0 86.9 86.9 87.0 86.9 87.0
Book reviews 77.2 76.4 77.2 75.8 83.5 78.7 83.7 89.9 90.7 91.0 90.7 91.1 90.4 91.0
Music reviews 76.6 76.5 79.6 70.9 82.1 76.6 82.7 87.8 89.2 89.2 89.0 89.1 88.8 89.5

Table 5: Recall comparison for different types of datasets.


Word2vec (%) BERT (%)
Datasets
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S
Sentiment140 (10%) 74.2 80.6 80.0 79.4 78.9 80.0 80.5 84.7 84.1 84.2 84.1 84.1 84.2 83.9
Tweets Airline 83.0 86.3 88.2 88.1 87.4 88.8 87.7 91.9 92.9 92.8 93.0 92.8 93.1 92.5
Tweets SemEval 80.1 84.1 87.4 88.3 86.4 85.2 85.0 91.6 92.2 92.1 92.3 91.6 91.8 92.1
IMDb movie reviews (1) 78.9 89.2 90.0 88.9 90.1 90.1 90.1 93.2 93.6 93.3 93.1 93.6 93.4 93.4
IMDb movie reviews (2) 82.9 86.9 85.4 87.6 89.6 89.6 89.3 90.6 90.9 90.7 90.9 90.8 90.9 90.8
Cornell movie reviews 67.1 73.6 75.1 71.1 77.4 75.7 75.6 84.5 87.2 86.6 86.6 86.8 86.7 87.1
Book reviews 72.6 78.0 78.2 76.3 83.4 78.2 83.4 90.7 90.8 90.8 90.5 91.0 90.9 91.0
Music reviews 75.7 75.8 79.4 70.8 80.8 76.5 82.8 87.7 88.5 88.8 88.2 88.3 87.9 89.2

Table 6: Precision comparison for different types of datasets.


Word2vec (%) BERT (%)
Datasets
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S
Sentiment140 (10%) 74.1 78.6 81.1 81.9 82.0 81.6 79.1 79.7 83.9 83.9 83.9 83.9 82.8 84.0
Tweets Airline 81.0 87.6 85.7 88.1 88.0 88.4 88.0 92.3 92.9 92.8 92.7 92.8 92.8 93.1
Tweets SemEval 81.1 82.2 83.0 78.3 83.6 83.7 84.1 90.7 91.7 91.5 91.5 91.9 91.5 91.5
IMDb movie reviews (1) 79 85.6 86.7 91.0 90.2 89.9 90.5 93.4 93.2 93.5 93.7 93.2 93.4 93.5
IMDb movie reviews (2) 82.7 87.9 84.8 91.5 89.3 89.2 89.6 90.3 90.6 90.6 90.5 90.7 90.4 90.5
Cornell movie reviews 69.8 70.8 78.4 82.0 74.8 77.3 76.3 86.3 86.9 87.5 87.4 87.4 87.1 86.9
Book reviews 88.3 74.8 76.3 76.3 84.0 79.9 84.0 89.0 90.7 91.2 90.9 91.4 89.8 91.1
Music reviews 79.7 81.4 80.1 76.7 84.5 77.2 82.7 88.1 90.2 89.9 90.3 90.3 90.1 90.0

Table 7: F-score comparison for different types of datasets.


Word2vec (%) BERT (%)
Datasets
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S
Sentiment140 (10%) 74.1 79.5 80.5 80.4 80.3 80.8 79.8 81.8 84.0 84.1 84.0 84.0 83.3 83.9
Tweets Airline 82.0 86.9 86.9 87.9 87.6 88.5 87.9 92.0 92.9 92.8 92.8 92.8 92.9 92.8
Tweets SemEval 80.6 83.1 84.9 82.8 84.9 84.4 84.5 91.1 91.9 91.8 91.9 91.8 91.6 91.8
IMDb movie reviews (1) 78.9 87.3 88.3 89.8 90.1 90.0 90.3 93.3 93.4 93.4 93.4 93.4 93.4 93.4
IMDb movie reviews (2) 82.8 87.4 85.0 89.5 89.4 89.4 89.5 90.4 90.7 90.6 90.7 90.7 90.6 90.7
Cornell movie reviews 68.3 71.5 76.6 75.1 76.0 76.5 75.9 85.4 87.0 87.0 87.0 87.1 86.9 87.0
Book reviews 79.5 75.9 76.7 75.6 83.5 79.0 83.7 89.8 90.7 91.0 90.7 91.1 90.3 91.0
Music reviews 77.3 77.3 79.6 72.5 82.3 76.7 82.7 87.8 89.3 89.3 89.1 89.2 88.9 89.6

produced better results than single models using either single models. Using BERT, the results have also improved
Word2vec or BERT. With the use of Word2vec, the results of although by a smaller amount since these models have
accuracy from hybrid models are higher than the ones from reached a relatively high accuracy, mostly more than 90%.
Complexity 9

Table 8: AUC comparison for different types of datasets.


Word2vec (%) BERT (%)
Datasets
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S
Sentiment140 (10%) 74.2 79.7 80.3 80.1 79.9 80.6 79.9 83.0 84.0 84.1 84.0 84.1 83.8 84.0
Tweets Airline 82.3 86.8 87.1 88.0 87.5 88.6 87.8 92.1 92.9 92.8 92.9 92.8 92.9 92.8
Tweets SemEval 80.5 84.3 86.2 84.6 86.0 85.5 85.6 91.2 92.0 91.8 92.0 91.8 91.7 91.8
IMDb movie reviews (1) 78.9 87.6 88.5 89.7 90.0 90.0 90.3 93.3 93.4 93.4 93.4 93.4 93.4 93.4
IMDb movie reviews (2) 82.9 87.3 85.1 89.2 89.4 89.4 89.5 90.5 90.7 90.7 90.7 90.7 90.7 90.7
Cornell movie reviews 67.8 72.3 76.1 73.0 76.4 76.2 75.9 85.3 87.1 87.0 86.9 87.0 86.9 87.0
Book reviews 79.0 76.4 77.2 75.8 83.5 78.7 83.7 90.0 90.8 91.0 90.7 91.2 90.5 91.1
Music reviews 77.2 76.5 79.6 70.9 82.1 76.6 82.7 87.9 89.3 89.3 89.2 89.2 88.9 89.6

Accuracy
95
90
85
80
(%)
75
70
65
60
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S
Word2vec (%) BERT (%)

Sentiment 140 (10%) Tweets Airline Tweets Semeval


IMDB Movie Reviews (1) IMDB Movie Reviews (2) Cornell Movie Reviews
Book Reviews Music Reviews
Figure 4: Accuracy values of deep learning models with Word2vec and BERT for different datasets.

Recall
95
90
85
80
(%)
75
70
65
60
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S
Word2vec (%) BERT (%)

Sentiment 140 (10%) Tweets Airline Tweets Semeval


IMDB Movie Reviews (1) IMDB Movie Reviews (2) Cornell Movie Reviews
Book Reviews Music Reviews
Figure 5: Recall values of deep learning models with Word2vec and BERT for different datasets.

The text in a review is normally longer than the text in a from only 1 to 50 words. Besides, the length of a tweet ranges
tweet, which suggests that LCNN-SVN performs better than from 1 to 40 words; however, the distribution of sample
other hybrid models on longer textual sample (Table 4). In length on Sentiment140 dataset is right skewed. It is ob-
selected datasets, when examining the distribution of the served that the results on two datasets, Sentiment140 and
textual length of samples, the length of review ranges from 1 Cornell movie reviews, are lower than those of the remaining
to 800 words. However, the Cornell movie reviews range datasets.
10 Complexity

Precision
95
90
85
80
(%)
75
70
65
60
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S
Word2vec (%) BERT (%)

Sentiment 140 (10%) Tweets Airline Tweets Semeval


IMDB Movie Reviews (1) IMDB Movie Reviews (2) Cornell Movie Reviews
Book Reviews Music Reviews
Figure 6: Precision values of deep learning models with Word2vec and BERT for different datasets.

F-Score
95
90
85
80
(%)
75
70
65
60
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S

Word2vec (%) BERT (%)

Sentiment 140 (10%) Tweets Airline Tweets Semeval


IMDB Movie Reviews (1) IMDB Movie Reviews (2) Cornell Movie Reviews
Book Reviews Music Reviews
Figure 7: F-score values of deep learning models with Word2vec and BERT for different datasets.

AUC
95
90
85

(%) 80
75
70
65
60
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S

Word2vec (%) BERT (%)

Sentiment 140 (10%) Tweets Airline Tweets Semeval


IMDB Movie Reviews (1) IMDB Movie Reviews (2) Cornell Movie Reviews
Book Reviews Music Reviews
Figure 8: AUC values of deep learning models with Word2vec and BERT for different datasets.
Complexity 11

Table 9: A comparison based on the proposed models and state-of-the-art approaches on datasets.
Study Model Dataset Accuracy (%)
Kim and Jeong [18] CNN Cornell movie reviews 81
Maulana et al. [75] SVM-IG Cornell movie reviews 85.65
Proposed hybrid model LCNN-SVM Cornell movie reviews 87
Jnoub et al. [19] SNN/CNN IMDb 87/81
McCann et al. [20] Char + CoVe-LSTM IMDb 92.1
Tang et al. [76] L-GRNN/Conv-GRNN IMDb 45.3/42.5
Maltoudoglou et al. [49] BERT IMDb 92.28
Yang et al. [47] XLNET IMDb 96.21
Proposed hybrid model CNN-LSTM IMDb 93.4
Baziotis et al. [77] Bi-LSTM + attention Tweets SemEval 67.7 (F1)
Cliché [78] LSTM-CNN Tweets SemEval 68.5 (F1)
Proposed hybrid model CNN-LSTM Tweets SemEval 91.9 (F1)
Abid et al. [12] Bi-LSTM/CNN Sentiment140 87.21/72.42
Han et al. [79] FK-SVM Sentiment140 87.2
Proposed hybrid model LSTM-CNN Sentiment140 84.1
Rane and Kumar [80] AdaBoost Tweets Airline 84.5
Duan et al. [81] SVM and Naive Bayes Tweets Airline 80
Monika et al. [17] LSTM Tweets Airline 80
Proposed hybrid model CLSTM-SVM Tweets Airline 92.9
Blitzer et al. [82] SCL-MI Music reviews/book reviews 79.7
Uribe [83] Logistic/SVM Music reviews/book reviews 87/89
Proposed hybrid model LSTM-CNN/LC-S Music reviews/book reviews 91.1/89.5

Some other studies performing sentiment analysis by complexity. Since time is one of the most valuable resources
using a single dataset of tweets or reviews are presented in and the most taken into account when evaluating the per-
[29, 33, 34, 37–39, 73, 74]. Note that the hybrid models formance of algorithms, we include the analysis of the
provide much improved results in terms of processing time computational time of the models involved in the com-
and accuracy. In addition, the overall accuracy of these parative study, as this is a reflection of the time complexity.
hybrid models was given with eight different types of Table 10 contains the time processing required for all
datasets, which give an objective view of overall accuracy. datasets involved in the experiments. Processing time is
Among the state-of-the-art approaches shown in Table 9, calculated for the entire process of training and testing
most of our hybrid models proposal got higher accuracy models using Word2vec and BERT. It includes time for data
results on six datasets. On Sentiment140, however, Han et al. division and time to create the classification model (initialize
[79] and Abid et al. [12] achieved a better accuracy of around the number of layers of the neural network, the number of
87%. The XLNet method for sentiment analysis with IMDb nodes per layer, etc.) but does not include the time used to
dataset, performed by Yang et al. [47], resulted in 96.21% display the classification results. When using the hybrid
accuracy. On the other hand, Akhtar et al. [37] tested a models with the BERT technique for feature extraction, the
hybrid model of combined CNN and SVM on both tweet accuracy generally is higher than with Word2vec, but the
and review datasets; however, the results showed a lower processing time is longer.
accuracy in comparison to hybrid methods, which were only In general, the hybrid methods provide better results
tested on a single type of dataset (58.62% accuracy on tweet than single deep learning models. Most hybrid networks
dataset and 77.16% accuracy with review dataset). The provide higher (or equal) scores in all datasets. Moreover,
comparison details with the state-of-the-art approaches are from the good result of Maltoudoglou et al. [49] (Table 9), we
shown in Table 9. It includes the authors’ names, methods, saw that the feature extraction plays an important role in
datasets, and accuracy (or F1 for some studies that only sentiment classification. We also discussed the importance
provide the F1 measure). of feature extraction in [1], where TF-IDF and word em-
In addition to the evaluation of the reliability of the bedding techniques for feature extraction were analysed.
models, it is also important to evaluate the performance of These improved results are high and stable at the expense of
the algorithms in terms of resource utilization. There is very some increase in processing time, as shown in Table 10. The
little work evaluating the computational complexity of deep table shows that the hybrid model required longer com-
learning models although there is some proposal [84] that putational time than the single models, because hybrid
considers some factors, such as the number of layers, the size models are complex and feature many more parameters than
of the input matrix, and other factors depending on the single models. While the computational times are longer,
specific algorithm. In CNN, the number and size of con- they do not preclude analysis of the trade-offs between
volution kernels and the number of output channels of each processing time and accuracy of results.
layer are considered. In view of this, it is clear that the higher Our aim is to build a hybrid deep learning model for
reliability of hybrid models comes at the cost of higher sentiment analysis that works well on various datasets of
12

Table 10: Time processing for experiments using a Google Colab Pro by dataset and model used.
Word2vec BERT
Datasets
SVM CNN LSTM C-L L-C CL-S LC-S SVM CNN LSTM C-L L-C CL-S LC-S
Sentiment140 (10%) 1h08m32 15m44 24m19 26m00 26m50 26m13 27m41 5h13m30 5h52m25 5h22m12 5h59m58 6h2m08 6h0m34 6h7m01
Tweets Airline 0m35 0m40 1m20 1m34 1m28 2m56 2m52 9m55 0h11m42 0h9m58 0h11m52 0h10m42 0h13m00 0h11m14
Tweets SemEval 1m18 1m32 2m04 2m30 2m20 3m55 3m06 8m46 0h9m46 0h8m40 0h10m12 0h10m07 0h14m28 0h37m48
IMDb movie reviews (1) 1h16m26 1h24m19 1h43m52 1h40m41 1h00m03 1h55m12 2h03m57 6h45m15 6h25m50 6h22m50 6h27m00 6h26m40 7h10m15 6h37m40
IMDb movie reviews (2) 39m42 48m04 0h54m03 58m27 1h03m23 1h0m42 1h05m15 3h24m20 3h11m00 3h9m30 3h11m40 3h11m25 3h36m15 3h28m05
Cornell movie reviews 2m16 2m46 3m42 4m49 4m19 6m04 5m45 16m12 0h15m26 0h14m10 0h15m55 0h15m49 0h24m03 0h19m41
Book reviews 2m23 5m31 6m46 9m40 8m35 10m16 9m21 32m22 0h33m11 0h32m37 0h33m23 0h33m49 0h33m59 0h58m18
Music reviews 1m43 2m41 4m36 5m53 5m21 5m21 5m45 18m23 0h18m44 0h18m17 0h18m55 0h19m03 0h20m14 1h5m57
Complexity
Complexity 13

domains. However, when building the classification models, Data Availability


there are many parameters that must be defined before, so
they can be suitable for a given dataset but not for others. The datasets used to support the findings of this study are
Therefore, the results obtained are positive and highly re- available from the direct link in the dataset citations.
liable because they have been evaluated on many datasets
with different topics. Finally, general summaries of the re- Disclosure
sults achieved in the experiments referenced earlier are
discussed as follows: The funders had no role in the design of the study; in the
collection, analyses, or interpretation of data; in the writing
(i) The hybrid models increased the accuracy for
of the manuscript; or in the decision to publish the results.
sentiment analysis compared with a single model
performance on all types of datasets, although the
computation time of SVM models is longer. Conflicts of Interest
(ii) The combination helped to take advantage of the The authors declare no conflicts of interest.
strengths of CNN, LSTM, and SVM, where CNN
has the capability to extract characteristics, LSTM
has capability to store past information at the Acknowledgments
state nodes (cell state), and SVM has capability to This work was supported by the Spanish Government and
classify. European (Fondo Europeo de Desarrollo Regional) FEDER
(iii) Using SVM as the classification method improved funds, project InEDGEMobility: Movilidad inteligente y
the results of both L-CNN and C-LSTM. SVM is sostenible soportada por Sistemas Multi-agentes y Edge
effective in multidimensional data stratification and Computing (RTI2018-095390-B-C32).
helps minimize local minima of neural networks.
References
6. Conclusions
[1] N. C. Dang, M. N. Moreno-Garcı́a, and F. De la Prieta,
In this paper, we proposed the use of hybrid deep learning “Sentiment analysis based on deep learning: a comparative
models for sentiment analysis from social network data. study,” Electronics, vol. 9, no. 3, p. 483, 2020.
We tested the performance of mixing SVM, CNN, and [2] M. J. S. Keenan, Advanced Positioning, Flow, and Sentiment
LSTM, using two-word embedding techniques, Word2vec Analysis in Commodity Markets: Bridging Fundamental and
and BERT, on eight textual datasets of tweets and reviews. Technical Analysis, Wiley, Hoboken, NJ, USA, 2nd edition,
Afterwards, we compared four generated hybrid models 2018.
[3] S. Sohangir, D. Wang, A. Pomeranets, and T. M. Khoshgoftaar,
with single models. These experiments are conducted to
“Big data: deep learning for financial sentiment analysis,”
understand the adaptability of hybrid models, whether Journal of Big Data, vol. 5, no. 1, p. 3, 2018.
hybrid approaches can adapt in a wide range of dataset [4] H. Jangid, S. Singhal, R. R. Shah, and R. Zimmermann,
types and sizes. We studied the influence of different types “Aspect-based financial sentiment analysis using deep
of datasets, feature extraction techniques, and deep learning,” in Companion Proceedongs of the Web Conference
learning models on reliability of sentiment polarity 2018, International World Wide Web Conferences Steering
analysis. Committee, Lyon, France, April 2018.
Our experiments reveal that the reliability of hybrid [5] G. Wang, G. Yu, and X. Shen, “The effect of online investor
models outperformed among all tested models for sentiment sentiment on stock movements: an LSTM approach,” Com-
polarity analysis. Combining deep learning models with the plexity, vol. 2020, Article ID 4754025, 11 pages, 2020.
[6] R. Satapathy, E. Cambria, and A. Hussain, Sentiment Analysis
SVM technique yields better results than using an individual
in the Bio-Medical Domain, Springer Interntional Publishing
model for performing sentiment analysis. In most of the AG, Basel, Switzerland, 2017.
tested datasets, the reliability of hybrid models using SVM is [7] A. Rajput, “Natural language processing, sentiment analysis,
higher than that of the ones not using it; however, the and clinical analytics,” in Innovation in Health Informatics,
computational time is much longer for the ones with SVM. pp. 79–97, Academic Press, Cambridge, MA, USA, 2020.
We also observed that the effectiveness of the algorithms [8] V. Malik and A. Kumar, “Sentiment analysis of twitter data
depends largely on the characteristics and quality of the using Naive Bayes algorithm,” International Journal on Recent
datasets. and Innovation Trends in Computing and Communication,
We are aware that the context of the dataset has a large vol. 6, no. 4, pp. 120–125, 2018.
impact on the choice of sentiment analysis models. We [9] P. Vateekul and T. Koomsubha, “A study of sentiment
intend to study the performance of hybrid approaches for analysis using deep learning techniques on Thai Twitter data,”
in 2016 13th International Joint Conference on Computer
sentiment analysis on hybrid datasets and multiple or hybrid
Science and Software Engineering (JCSSE), IEEE, Khon Kaen,
contexts in order to gain deeper insight in a specific topic, Thailand, July 2016.
such as business, marketing, or medicine. Its application [10] A. C. Pandey, D. S. Rajpoot, and M. Saraswat, “Twitter
derives from associating sentiments to relevant context in sentiment analysis using hybrid cuckoo search method,”
order to provide detailed personal feedback and recom- Information Processing & Management, vol. 53, no. 4,
mendation for users. pp. 764–779, 2017.
14 Complexity

[11] A. S. M. Alharbi and E. de Doncker, “Twitter sentiment [27] M. Elleuch, R. Maalej, and M. Kherallah, “A new design
analysis with a deep neural network: an enhanced approach based-SVM of the CNN classifier architecture with dropout
using user behavioral information,” Cognitive Systems Re- for offline Arabic handwritten recognition,” Procedia Com-
search, vol. 54, pp. 50–61, 2019. puter Science, vol. 80, pp. 1712–1723, 2016.
[12] F. Abid, M. Alam, M. Yasir, and C. Li, “Sentiment analysis [28] Y. Tang, “Deep learning using linear support vector ma-
through recurrent variants latterly on convolutional neural chines,” 2013, [Link]
network of Twitter,” Future Generation Computer Systems, [29] T. Chen, R. Xu, Y. He, and X. Wang, “Improving sentiment
vol. 95, pp. 292–308, 2019. analysis via sentence type classification using BiLSTM-CRF
[13] A. M. Ramadhani and H. S. Goo, “Twitter sentiment analysis and CNN,” Expert Systems with Applications, vol. 72,
using deep learning methods,” in 2017 7th International pp. 221–230, 2017.
Annual Engineering Seminar (InAES), IEEE, Yogyakarta, [30] A. U. Rehman, A. K. Malik, B. Raza, and W. Ali, “A hybrid
Indonesia, August 2017. CNN-LSTM model for improving accuracy of movie reviews
[14] A. M. Khattak, R. Batool, F. A. Satti et al., “Tweets classifi- sentiment analysis,” Multimedia Tools and Applications,
cation and sentiment analysis for personalized tweets rec- vol. 78, no. 18, pp. 26597–26613, 2019.
ommendation,” Complexity, vol. 2020, Article ID 8892552, [31] Q.-H. Vo, H.-T. Nguyen, B. Le, and M.-L. Nguyen, “Multi-
11 pages, 2020. channel LSTM-CNN model for Vietnamese sentiment anal-
[15] A. Hassan and A. Mahmood, “Deep learning approach for ysis,” in 2017 9th international conference on knowledge and
sentiment analysis of short texts,” in Third International systems engineering (KSE), IEEE, Hue, Vietnam, October
Conference on Control, Automation and Robotics (ICCAR), 2017.
IEEE, Nagoya, Japan, April 2017. [32] C. A. Martı́n, J. M. Torres, R. M. Aguilar, and S. Diaz, “Using
[16] J. Qian, Z. Niu, and C. Shi, “Sentiment analysis model on deep learning to predict sentiments: case study in tourism,”
weather related tweets with deep neural network,” in Pro- Complexity, vol. 2018, Article ID 7408431, 9 pages, 2018.
ceedings of the 2018 10th International Conference on Machine [33] K. Elshakankery and M. F. Ahmed, “HILATSA: a hybrid
Learning and Computing, ACM, Zhuhai, China, February Incremental learning approach for Arabic tweets sentiment
2018. analysis,” Egyptian Informatics Journal, vol. 20, no. 3,
[17] R. Monika, S. Deivalakshmi, and B. Janet, “Sentiment analysis pp. 163–171, 2019.
of US airlines tweets using LSTM/RNN,” in 2019 IEEE 9th [34] S. J. Putra, I. Khalil, M. N. Gunawan, R. I. Amin, and
International Conference on Advanced Computing (IACC), T. Sutabri, “A hybrid model for social media sentiment
IEEE, Tiruchirappalli, India, December 2019. analysis for Indonesian text,” in Proceedings of the 20th
[18] H. Kim and Y.-S. Jeong, “Sentiment classification using International Conference on Information Integration and
convolutional neural networks,” Applied Sciences, vol. 9, Web-Based Applications & Services, Yogyakarta, Indonesia,
no. 11, p. 2347, 2019. November 2018.
[19] N. Jnoub, F. Al Machot, and W. Klas, “A domain-independent [35] P. Astya, “Sentiment analysis: approaches and open issues,” in
classification model for sentiment analysis using neural 2017 International Conference on Computing, Communication
models,” Applied Sciences, vol. 10, no. 18, p. 6221, 2020. and Automation (ICCCA), IEEE, Greater Noida, India, May
[20] B. McCann, J. Bradbury, C. Xiong, and R. Socher, “Learned in 2017.
translation: contextualized word vectors,” in Proceedings of [36] H. Srinidhi, G. Siddesh, and K. Srinivasa, “A hybrid model
the 31st International Conference on Neural Information using MaLSTM based on recurrent neural networks with
Processing Systems, Long Beach, CA, USA, December 2017. support vector machines for sentiment analysis,” Engineering
[21] H. Mhaskar, Q. Liao, and T. Poggio, “When and why are deep and Applied Science Research, vol. 47, no. 3, pp. 232–240, 2020.
networks better than shallow ones?” in Proceedings of the 2017 [37] M. S. Akhtar, A. Kumar, A. Ekbal, and P. Bhattacharyya, “A
AAAI Conference on Artificial Intelligence, San Francisco, CA, hybrid deep learning architecture for sentiment analysis,” in
USA, February 2017. Proceedings of COLING 2016, the 26th International Con-
[22] A. Schindler, T. Lidy, and A. Rauber, “Comparing shallow ference on Computational Linguistics: Technical Papers, Osaka,
versus deep neural network architectures for automatic music Japan, December 2016.
genre classification,” in Proceedings of the 9th Forum Media [38] S. Al-Azani and E.-S. M. El-Alfy, “Hybrid deep learning for
Technology (FMT2016), FMT, Poelten, Austria, 2016. sentiment polarity determination of Arabic microblogs,” in
[23] I. Gupta and N. Joshi, “Enhanced twitter sentiment analysis International Conference on Neural Information Processing,
using hybrid approach and by accounting local contextual Springer, Guangzhou, China, November 2017.
semantic,” Journal of Intelligent Systems, vol. 29, no. 1, [39] G. Liu, X. Xu, B. Deng, S. Chen, and L. Li, “A hybrid method
pp. 1611–1625, 2019. for bilingual text sentiment classification based on deep
[24] J. F. Sánchez-Rada and C. A. Iglesias, “CRANK: a hybrid learning,” in Proceedings of the 2016 17th IEEE/ACIS Inter-
model for user and content sentiment classification using national Conference on Software Engineering, Artificial In-
social context and community detection,” Applied Sciences, telligence, Networking and Parallel/Distributed Computing
vol. 10, no. 5, p. 1662, 2020. (SNPD), IEEE, Shanghai, China, May 2016.
[25] R. Alfrjani, T. Osman, and G. Cosma, “A hybrid semantic [40] Q. Zhang, Z. Zhang, M. Yang, and L. Zhu, “Exploring co-
knowledgebase-machine learning approach for opinion evolution of emotional contagion and behavior for microblog
mining,” Data & Knowledge Engineering, vol. 121, pp. 88–108, sentiment analysis: a deep learning architecture,” Complexity,
2019. vol. 2021, Article ID 6630811, 10 pages, 2021.
[26] D.-X. Xue, R. Zhang, H. Feng, and Y.-L. Wang, “CNN-SVM [41] H. Kaur, S. U. Ahsaan, B. Alankar, and V. Chang, “A proposed
for microvascular morphological type recognition with data sentiment analysis deep learning algorithm for analyzing
augmentation,” Journal of Medical and Biological Engineering, COVID-19 tweets,” Information Systems Frontiers, pp. 1–13,
vol. 36, no. 6, pp. 755–764, 2016. 2021.
Complexity 15

[42] Z. Kastrati, L. Ahmedi, A. Kurti et al., “A deep learning [59] “Multi-domain sentiment dataset,” Available from: (accessed
sentiment analyser for social media comments in low-re- on 10 December 2020), [Link]
source languages,” Electronics, vol. 10, no. 10, p. 1133, 2021. datasets/sentiment/.
[43] A. H. Ombabi, W. Ouarda, and A. M. Alimi, “Mining, “deep [60] Y. Wan and Q. Gao, “An ensemble sentiment classification
learning CNN–LSTM framework for Arabic sentiment system of twitter data for airline services analysis,” in Pro-
analysis using textual information shared in social networks,” ceedings of the 2015 IEEE International Conference On Data
Social Network Analysis and Mining, vol. 10, no. 1, pp. 1–13, Mining Workshop (ICDMW), IEEE, Atlantic City, NJ, USA,
2020. November 2015.
[44] G. Yoo and J. Nam, “A hybrid approach to sentiment analysis [61] L. Zhang, S. Wang, and B. Liu, “Deep learning for sentiment
enhanced by sentiment lexicons and polarity shifting devices,” analysis: a survey,” WIREs Data Mining and Knowledge
in The 13th Workshop on Asian Language Resources, Miyazaki, Discovery, vol. 8, no. 4, Article ID e1253, 2018.
Japan, May 2018. [62] T. Mikolov, K. Chen, G. S. Corrado, and J. A. Dean,
[45] Y. Wang, M. Wang, and W. Xu, “A sentiment-enhanced “Computing numeric representations of words in a high-
hybrid recommender system for movie recommendation: a dimensional space,” Google Patents, 2015.
big data analytics framework,” Wireless Communications and [63] N. Banić and N. Elezović, “TVOR: finding discrete total
Mobile Computing, vol. 2018, Article ID 8263704, 9 pages, variation outliers among histograms,” IEEE Access, vol. 9,
2018. pp. 1807–1832, 2020.
[46] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: pre- [64] B. Jang, M. Kim, G. Harerimana, S.-U. Kang, and J. W. Kim,
training of deep bidirectional transformers for language “Bi-LSTM model to increase accuracy in text classification:
understanding,” 2018, [Link] combining Word2vec CNN and attention mechanism,” Ap-
[47] Z. Yang, Z. Dai, Y. Yang et al., “Xlnet: generalized autore- plied Sciences, vol. 10, no. 17, p. 5841, 2020.
gressive pretraining for language understanding,” 2019, [65] A. Jacovi, O. S. Shalom, and Y. Goldberg, “Understanding
[Link] convolutional neural networks for text classification,” 2018,
[48] A. Tela, A. Woubie, and V. Hautamaki, “Transferring [Link]
[66] H. T. Nguyen and M. Le Nguyen, “An ensemble method with
monolingual model to low-resource language: the case of
sentiment features and clustering support,” Neurocomputing,
tigrinya,” 2020, [Link]
[49] L. Maltoudoglou, A. Paisios, and H. Papadopoulos, “BERT- vol. 370, pp. 155–165, 2019.
[67] R. Yamashita, M. Nishio, R. K. G. Do, and K. Togashi,
based conformal predictor for sentiment analysis,” in
“Convolutional neural networks: an overview and application
Proceedings of the 2020 Conformal and Probabilistic Pre-
in radiology,” Insights into imaging, vol. 9, no. 4, pp. 611–629,
diction and Applications, PMLR, Verona, Italy, September
2018.
2020.
[68] S. Hochreiter and J. Schmidhuber, “LSTM can solve hard long
[50] X.-R. Gong, J.-X. Jin, and T. Zhang, “Sentiment analysis using
time lag problems,” in Proceedings of the 9th International
autoregressive language modeling and broad learning sys-
Conference on Neural Information Processing Systems, Denver,
tem,” in 2019 IEEE International Conference on Bioinformatics
CO, USA, December 1997.
and Biomedicine (BIBM), IEEE, San Diego, CA, USA, No- [69] M. M. Adankon, M. Cheriet, and A. Biem, “Semisupervised
vember 2019. least squares support vector machine,” IEEE Transactions on
[51] B. Myagmar, J. Li, and S. Kimura, “Cross-domain sentiment Neural Networks, vol. 20, no. 12, pp. 1858–1870, 2009.
classification with bidirectional contextualized transformer [70] “Making the most of your Colab subscription,” Available
language models,” IEEE Access, vol. 7, pp. 163219–163230, from: (accessed on 22 January 2021), [Link]
2019. [Link]/notebooks/[Link].
[52] S. Kumar, M. Gahalawat, P. P. Roy, D. P. Dogra, and [71] “Keras: The Python deep learning API,” Available from:
B.-G. Kim, “Exploring impact of age and gender on sentiment (accessed on 10 December 2020, [Link]
analysis using machine learning,” Electronics, vol. 9, no. 2, [72] “TensorFlow,” Available from: (accessed on 10 December
p. 374, 2020. 2020), [Link]
[53] “Sentiment140 - a twitter sentiment analysis tool,” Available [73] K. Ghasedi and H. Huang, “Sentiment analysis via deep
from: (accessed on 10 December 2020), [Link] hybrid textual-crowd learning model,” in Proceedings of the
[Link]/site-functionality. Thirty-Second AAAI Conference on Artificial Intelligence
[54] “Twitter US Airline Sentiment,” Available from: (accessed on (AAAI 2018), New Orleans, Louisiana, February 2018.
10 December 2020), [Link] [74] M. U. Salur and I. Aydin, “A novel hybrid deep learning
twitter-airline-sentiment. model for sentiment classification,” IEEE Access, vol. 8,
[55] “International Workshop on Semantic Evaluation 2017, pp. 58080–58093, 2020.
Available from: (accessed on 10 December 2020), [Link] [75] R. Maulana, P. A. Rahayuningsih, W. Irmayani, D. Saputra,
[Link]/semeval2017/. and W. E. Jayanti, “Improved accuracy of sentiment analysis
[56] “Large movie review dataset,” Available from: (accessed on 10 movie review using support vector machine based informa-
December 2020), [Link] tion gain,” in Proceedings of the International Conference on
sentiment/. Advanced Information Scientific Development (ICAISD), IOP
[57] “Bag of words meets bags of popcorn,” Available from: Publishing, West Java, Indonesia, August 2020.
(accessed on 10 December 2020), [Link] [76] D. Tang, B. Qin, and T. Liu, “Document modeling with gated
word2vec-nlp-tutorial/data?select=[Link]. recurrent neural network for sentiment classification,” in
[58] “Cornell CIS computer science,” Available from: (accessed on Proceedings of the 2015 Conference On Empirical Methods In
10 December 2020), [Link] Natural Language Processing, Lisbon, Portugal, September
movie-review-data/. 2015.
16 Complexity

[77] C. Baziotis, N. Pelekis, and C. Doulkeridis, “Datastories at


semeval-2017 task 4: deep lstm with attention for message-
level and topic-based sentiment analysis,” in Proceedings of the
11th International Workshop On Semantic Evaluation
(SemEval-2017), Vancouver, Canada, August 2017.
[78] M. Cliche, “Bb_twtr at semeval-2017 task 4: twitter sentiment
analysis with cnns and lstms,” 2017, [Link]
06125.
[79] K.-X. Han, W. Chien, C.-C. Chiu, and Y.-T. Cheng, “Ap-
plication of support vector machine (SVM) in the sentiment
analysis of twitter DataSet,” Applied Sciences, vol. 10, no. 3,
p. 1125, 2020.
[80] A. Rane and A. Kumar, “Sentiment classification system of
Twitter data for US airline service analysis,” in Proceedings of
the 2018 IEEE 42nd Annual Computer Software and Appli-
cations Conference (COMPSAC), IEEE, Tokyo, Japan, July
2018.
[81] X. Duan, T. Ji, and W. Qian, Twitter US Airline Recom-
mendation Prediction, p. cs229, Stanford University, Stanford,
CA, USA, 2016.
[82] J. Blitzer, M. Dredze, and F. Pereira, “Biographies, bollywood,
boom-boxes and blenders: domain adaptation for sentiment
classification,” in Proceedings of the Association for Compu-
tational Linguistics (ACL), Prague, Czech Republic, June 2007.
[83] D. Uribe, “Domain adaptation in sentiment classification,” in
2010 Ninth International Conference on Machine Learning
and Applications, IEEE, Washington, DC, USA, December
2010.
[84] H. Xie, M. Zhang, J. Ge, X. Dong, and H. Chen, “Learning air
traffic as images: a deep convolutional neural network for
airspace operation complexity evaluation,” Complexity,
vol. 2021, Article ID 6457246, 16 pages, 2021.

You might also like