F-RW-02
Research Plan
(Font-Times New Roman & size: 36)
ASPECT BASED SENTIMENT ANALYSIS USING MACHINE LEARNING
Contents Page Number
1. Proposed Topic of Research
2. Research Objective
3. Introduction
4. Literature Review
5. Proposed method to be used in solving the problem /study
6. Material and Methodology
7. References
1. Proposed Topic of Research:
Aspect based sentiment analysis using machine learning
2. Research Objective
Sentiment analysis or opinion mining is one of the major tasks of NLP(Natural Language
Processing). It has gained lot of attention in recent years. With the advancement of social
media, people are now more connected and are able to post their own content on various sites.
Sentiment analysis works on discovering opinions, classifying the attitude they convey, and
ultimately categorizing them division-wise. Major objective of this research is to analyse the
textual data to help businesses monitor brand and customer feedback for a product thereby
understanding the customer needs.
In this research work we are making use of machine learning algorithms in order to classify the
aspect of the entities as positive, negative and neutral. Machine learning strategies work by
training an algorithm with a training data set before applying it to the actual data set. After
classification, performance of the algorithms is computed using performance matrics.
3. Introduction
Sentiment analysis (SA) is one of the techniques in natural language processing (NLP) which
helps in understanding the sentiments, which may help the entrepreneurs to get information
about their customer views through different online mediums like social media, e-commerce
site reviews, surveys etc.
Sentiment analysis is the field that deals with the feelings, judgments as well as responses that
we get from the texts. This is widely used in various fields like data mining, web mining, and
social media analytics because sentiments are the most essential characteristics to judge the
human behaviour.
Sentiment analysis is done since every review we get does not directly tell us whether it is
“good” or a “bad” notion.
It can be performed for a particular document, sentence or aspect/feature of an entity. The
document level analysis is performed for certain sets of statements, sentence level analysis is
performed for detecting overall sentiment of a sentence and aspect level is performed for
determining the sentiment of a person for an aspect of an entity.
When we analyze sentiments through the texts or say product reviews, we need to analyse the
features that is been monitored in a positive, neutral, or negative way. That's where aspect based
sentiment analysis can help, for example in the given text "xyz phone does not have good sound
quality”, so the aspect-based classifier will determine that the sentence expresses a negative
opinion about the feature sound quality.
There are different ways of classifying the sentiments, the Machine learning approach and the
lexicon-based approach as shown in fig 1. The Machine Learning approach can further be
classified as supervised and unsupervised learning.
Supervised learning algorithm like SVM, Naive bayes, decision tree are used to classify the
reviews into positive and negative classes based on training data. The machine learning
approach is used to classify the data as positive or negative by using different algorithms.
While supervised learning can be defined as the process of learning from already known data to
generate initially a model and further predict target class for the particular data, unsupervised
learning can be defined as the process of learning from unlabeled to discriminate the provided
input data.
SA applications have been widely spread to nearly every domain, like social events, consumer
products and services, healthcare, political elections and financial services.
Fig 1. Classification of sentiment analysis
4. Literature Review
To perform sentiment analysis, Lexicon based approach can be performed and another one is
machine learning based approach. Both approaches have their own pros and cons. Aspect level
sentiment analysis analyzes the data by concentrating on features or aspects of the data.
Machine learning is also a mainstream sentiment classification method. These methods mainly
involve text representation and feature extraction, then training a sentiment classifier, such as
support vector machine (SVM) and logistic regression (LR).
The Aspect level sentiment analysis problem was first defined by Hu and Liu [1] ,they
introduced the distinction between explicit and implicit aspects and, since then, it has received
great attention. Sentiment analysis has been generally categorized at three levels. Document-
level [12], sentence-level [13] and aspect-level SA [14] to classify whether a whole document,
sentence (subjective or objective) or an aspect expresses a sentiment, i.e., positive, negative or
neutral.
A survey by Nazir et al. [2] identifies two main tasks for Aspect level sentiment analysis:
aspect extraction and aspect sentiment analysis.
Manek et al. [3] proposed the method of feature extraction based on gini index and
classification by support vector machine (SVM).
Hai et al. [4] proposed a new probabilistic supervised joint emotion model (SJSM), which could
not only identify semantic sentiments from the comment data but also infer the overall
sentiment of the comment data.
Singh et al. [5] used naive bayes, J48, BFTree and OneR four machine learning algorithms for
text sentiment analysis.
Huang et al. [6] proposed a multi-modal joint sentiment theme model. Based on the
introduction of user personality features and sensitive influence factors, the model uses latent
dirichlet allocation (LDA) model to analyze the hidden user sentiments and topic types in
Weibo text.
Huq et al. [7] used SVM and k-nearest neighbors (KNN) algorithms to analyze the sentiment of
twitter data.
Long et al. [8] used SVM to classify stock forum posts using additional samples containing
prior knowledge. Although the machine learning-based method can automatically extract
features, it often relies on manual feature selection. However, the deep learning-based approach
does not require manual intervention at all. It can automatically select and extract features
through the neural network structure and can learn from its own errors.
Jain and Dandannavar et al. [9] examined several steps for sentiment analysis on Twitter data
using machine learning algorithms. They also provided details of the proposed approach for
sentiment analysis. The approach collected data and then pre processed tweets using NLP-based
techniques. Afterward, feature extraction was performed to export sentiment-relevant features.
Finally, a model was trained using machine learning classifiers, such as naive Bayes classifiers,
support vector machine (SVM), and decision tree.
J. Zhou et al.[10] survey provides the most comprehensive overview of modern deep learning
techniques for Aspect based analysis. By collecting almost all of the standard ASC (aspect level
sentiment classification) datasets, including SemEval 2014, SemEval 2015, SemEval
2016, Twitter and other performed the analysis, compared and summarized the corresponding
algorithms.
[Link] et al.[11] proposed an efficient method for sentiment analysis by effectively
combining three procedures (a) creating the ontologies for extraction of semantic features (b)
Word2vec for conversion of processed corpus (c) convolutional neural network (CNN) for
opinion mining.
Lu et al.[17] employed a Support Vector Regression model to find the sentiment score for an
aspect. This allows the sentiment score to be modelled as a real number in the zero to five
interval.
Various algorithms have been used for sentiment classification SC, opinion extraction OE,
emotion detection ED, feature selection FS, aspect extraction AE and sentiment analysis SA. I
have categorised them in tabular form along with the data set used.
Approaches for sentiment Analysis
Algorithms
Year Task Data scope Data set/source
used
Markov
Bai X et Movie Reviews,
SC Blanket, SVM, IMDB
al(2010) News Articles
NB
Graph-Based
Camera Chinese Opinion
Yan et al(2010) SC approach, NB,
Reviews Analysis Domain
SVM
Weakly and Movie Reviews,
Yulan et IMDB,
SC semi supervised Multi-domain
al(2010) [Link]
classification sentiment
Digital
Multi-class
Chen et al(2011) SA Cameras, MP3
SVM
Reviews
[Link],
SVM, CHI Buyers’ posts
Teng et al(2011) SA [Link],
square web pages
[Link]
Kang et Restaurant
SC NB, SVM NA
al(2012)[25] Reviews
Alexandra et al Lexicon- Emotions
ED ISEAR, Emotinet
(2012)[26] Based, SVM corpus
SVM, K-
nearest
Lane et al(2012) media-analysis
SA neighbor,NB, MEDIA
[27] company
BN, DT, a Rule
learner
NEWS,
Reyes et al SEMANTIC,
FS CUSTOMER [Link]
(2012)[28] NB, SVM, DT
REVIEWS
Xianghua et al unsupervised,
SC social Reviews blog data set
(2012)[29] LDA
Martin et al
SC SVM,NB,C4.5 Film Reviews MCE Corpus
(2013) [30]
23656 mobile
Sun et al (2014)
AE Context Based phone users
[2]
review
Unsupervised,
Poria et al. Restaurant and SemEval 2014
AE rule-based
(2014) [32] product reviews dataset
approach
Schouten and
Restaurant and SemEval 2014
Frasincar (2016) AE Supervised
product reviews dataset
[14]
supervised,
Xu & Zang cell phone
AE Svm based
(2015) [31] review
approach
Manek et al FS, Gini index,
Movie Review IMDB
(2016)[3] SC SVM
Jain and
NB, Decision
Dandannavar et SC Tweets Twitter
Tree
al(2016)[9]
Naïve Bayes,
Singh et al product review,
SA J48, BFTree amazon, IMDB
(2017)[5] movie review
and OneR
Chatterji,S et al. Tweets, reviews
AE supervised twitter
(2017) [33] and news
Huq et al (2017)
SC KNN, SVM Tweets twitter
[7]
Jinzhan Feng deep 1.25 million
et. Al (2018) CNN & sequent mobile phone
[13] ial algorithm reviews
5. Proposed method to be used in solving the problem/study
Once the data is extracted from various data sets eg say Twitter API, which is publicly
available in NLTK resource we perform following steps-
1)Preprocessing is first step for which, with the help of an algorithm the noise is removed from
the raw data i.e removing all hyperlinks, hashtags, removing repeated letters, case conversions,
replacing all emoticons with their corresponding sentiments.
2 Classification Algorithms for Sentiment Analysis – initially I will be applying several
popular and commonly used classification algorithms such as the Multinomial Naïve Bayes
Algorithm, K-Nearest Neighbor Algorithm, SVM to identify sentiment polarity of the aspect
under consideration. After using the popular algorithms I will work and propose a new
algorithm which will have better performance.
3. Evaluation Metrics
For evaluation of performance of Aspect Based Sentiment Analysis model, the mostly used
evaluation metrics are accuracy, precision, recall and F1-measure. Accuracy is the proportion of
correctly predicted samples. But accuracy is insufficient measure if the dataset is unbalanced,
so precision, recall and F1-measure will be measured along with accuracy. The precision is
proportion of correct predictions among all positive label samples and recall is proportion of
correct predictions among all positive predictions. F1-measure is equally weighted average of
recall and precision.
6. Material and Methodology
The steps for performing aspect based sentiment analysis are shown in fig 2.
REVIEW COLLECTION PRE PROCESSING ASPECT EXTRACTION/
FEATURE SELECTION
SENTIMENT MAPPING/ POLARITY
CLASSIFICATION CALCULATED
Fig2 Steps for aspect based sentiment analysis
1. Review collection/Data collection
The data is the first thing that we need for sentiment analysis. The web is full of external
information from social media, news articles, product reviews, etc. And more and more
companies are making their datasets public. Following are the ways to collect relevant
information from different websites-
a) Web Scraping Tools- Web scraping tools, or web data extraction tools, are essential when it
comes to collecting external data.
b)APIs-These allow applications to communicate with another. To extract useful data from
websites or social media platforms, connect them with an API. Large companies like Facebook,
Twitter, and Instagram have their own APIs and allows to extract data from their platforms, to
gather comments from social media platforms about specific product features using aspect-
based analysis.
2. Aspect identification- Identification of the aspects specifies identification of words or
phrases which relates to the features of the review comments. For example, our product is
earphones; the important aspects of an earphone are say sound, size, design and cost.
[Link] of data-
An initial step in text and sentiment classification is pre-processing. This phase is one of the
most important phases in which cleaning of data and removal of stop words etc. happens to
improve the effectiveness of results.
• Vectorization phase- provides a record of data that will be required for classification of
reviews and a technique of vector space model is utilized for the same.
• POS(part of speech) tagging is part of speech tagging which permits to tag each word of data
to the POS i.e. verb, adverb, noun, pronoun, adjective etc.
• Stemming and lemmatization helps to reduce spatiality in the words. For example, the words
like ‘bright’, ‘brighter’, and ‘brightening’ are taken as one word ‘bright’.
• Stop word removal works on removing those words from data which do not affect the final
sentiment value of the data.
4. Classification- In order to classify the text into various classes we have three different
techniques i.e machine learning, lexicon based approach and hybrid. In my research work I will
be focussing on machine learning approaches.
Machine Learning
This approach, employes a machine-learning technique and diverse features to construct a
classifier that can identify text that expresses sentiment. Deep-learning methods are most
popular because they fit on data learning representations.
7. References
1. M. Hu and B. Liu, ``Mining and summarizing customer reviews,'' in Proc. ACM
SIGKDD Int. Conf. Knowl. Discovery Data Mining, 2004, pp. 168_177.
2. A. Nazir, Y. Rao, L. Wu, and L. Sun, ``Issues and challenges of aspect based sentiment
analysis: A comprehensive survey,'' IEEE Trans. Affect. Comput., early access, Jan. 30,
2020, doi: 10.1109/TAFFC.2020.2970399.
3. A. S. Manek, P. D. Shenoy, M. C. Mohan, and K. Venugopal, ``Aspect term extraction
for sentiment analysis in large movie reviews using Gini Index feature selection method
and SVM classifier,'' World Wide Web, vol. 20, no. 2, pp. 135_154, Mar. 2017.
4. Z. Hai, G. Cong, K. Chang, P. Cheng, and C. Miao, “Analyzing sentiments in one go:
A supervised joint topic modeling approach,'' IEEE Trans. Knowl. Data Eng., vol. 29,
no. 6, pp. 1172_1185, Jun. 2017.
5. J. Singh, G. Singh, and R. Singh, ``Optimization of sentiment analysis using machine
learning classifiers,'' Hum.-Centric Comput. Inf. Sci., vol. 7, no. 1, p. 32, 2017.
6. F. Huang, S. Zhang, J. Zhang, and G. Yu, “Multimodal learning for topic sentiment
analysis in microblogging,'' Neurocomputing, vol. 253, pp. 144_153, Aug. 2017.
7. M. R. Huq, A. Ali, and A. Rahman, “Sentiment analysis on Twitter data using KNN
and SVM,'' Int. J. Adv. Comput. Sci. Appl., vol. 8, no. 6, pp. 19_25, 2017.
8. W. Long, Y.-R. Tang, and Y.-J. Tian, ``Investor sentiment identification based on the
universum SVM,'' Neural Comput. Appl., vol. 30, no. 2, pp. 661_670, Jul. 2018.
9. A. P. Jain and P. Dandannavar, ``Application of machine learning techniques to
sentiment analysis,'' in Proc. 2nd Int. Conf. Appl. Theor. Comput. Commun. Technol.
(iCATccT), Jul. 2016, pp. 628_632.
10. Jie Zhou and Jimmy Xiangji Huang “Deep learning for aspect-level sentiment classification:
survey, vision, and challenges” IEEE Trans. IEEE Access, June 28, 2019,doi
10.1109/ACCESS.2019.2920075.
11. Ravindra Kumar, Husanbir Singh Pannu and Avleen Kaur Malhi “Aspect-based
sentiment analysis using deep networks and stochastic optimization” Neural Comput.
Appl (2017) 042009 doi:10.1088/1757-899X/263/4/042009.
12. A. Tripathy, A. Anand, and S.K. Rath, "Document-level Sentiment Classification using
Hybrid Machine Learning Approach," Knowledge Information Systems, vol. 53, no. 3,
pp. 805-831, 2017.
13. B. Liu, "Sentiment Analysis and Subjectivity," Handbook of Natural Language
Processing, pp. 627-666, 2010.
14. K. Schouten and F. Frasincar, "Survey on Aspect-Level Sentiment Analysis," IEEE
Trans. Knowledge and Data Eng., vol. 28, no. 3, pp. 813-830, 2016.
15. M. S. Akhtar, D. Gupta, A. Ekbal, and P. Bhattacharyya, "Feature Selection and
Ensemble Construction: A Two-Step Method for Aspect Based Sentiment Analysis,"
Knowledge-Based Systems, vol. 125, pp. 116-135, 2017.
16. C. Wu, F. Wu, S. Wu, Z. Yuan, and Y. Huang, "A Hybrid Unsupervised Method for
Aspect Term and Opinion Target Extraction," Knowledge-Based Systems, vol. 148, pp.
66-73, 2018.
17. B. Lu, M. Ott, C. Cardie, and B. K. Tsou, “Multi-Aspect Sentiment Analysis with Topic
Models,” in Proceedings of the 2011 IEEE 11th International Conference on Data
mining Workshops (ICDMW 2011). IEEE, 2011, pp. 81–88.
18. Poria, S., Cambria, E., & Gelbukh, A. (2016). “Aspect extraction for opinion mining
with a deep convolutional neural network” Knowledge-Based Systems, 108, 42–
49. doi:10.1016/[Link].2016.06.009
19. X. Chen, Y. Xue, H. Zhao, X. Lu, X. Hu, and Z. Ma, “A novel feature extraction
methodology for sentiment analysis of product reviews,” Neural Computing and
Applications, pp. 1–18, 2018.
20. Tang, D., Qin, B., & Liu, T. (2015a). Deep learning for sentiment analysis: Successful
approaches and future challenges. Wiley Interdisciplinary Reviews: Data Mining and
Knowledge Discovery, 5(6), 292–[Link]://[Link]/10.1002/widm.1171
21. Tang, D., Qin, B., & Liu, T. (2015b). Document Modeling with Gated Recurrent Neural
Network for Sentiment Classification. In Proceedings of the 2015 Conference on
Empirical Methods in Natural LanguageProcessing (pp. 1422–1432).
22. Tang, D., Qin, B., & Liu, T. (2016). Aspect Level Sentiment Classification with Deep
Memory Network. ArXiv Preprint ArXiv:1605.08900. Retrieved from
23. Xu, L., Lin, J., Wang, L., Yin, C., & Wang, J. “Deep Convolutional Neural Network
based Approach for Aspect-based Sentiment Analysis”,Advanced Science and
Technology Letters , 143(Ast), 199–204, 2017
24. Xu, L., Liu, J., Wang, L., & Yin, C ”Aspect Based Sentiment Analysis for Online
Reviews”. Advances in Computer Science and Ubiquitous Computing (Lecture No, Vol.
474). Springer Singapore, 2018.
25. Kang Hanhoon, Yoo Seong Joon, Han Dongil “Senti-lexicon and improved Naıve
Bayes algorithms for sentiment analysis of restaurant reviews” Expert Syst Appl
2012;39:6000–10.
26. Balahur Alexandra, Hermida Jesus M, Montoyo Andres “Detecting implicit expressions
of emotion in text: a comparative analysis.” Decis Support Syst 2012;53:742–53.
27. Lane Peter CR, Clarke Daoud, Hender Paul “On developing robust models for
favourability analysis: model choice, feature sets and imbalanced data.” Decis Support
Syst 2012;53:712–8.
28. Reyes Antonio, Rosso Paolo “Making objective decisions from subjective data:
detecting irony in customer reviews.” Decis Support Syst 2012;53:754–60.
29. Xianghua Fu, Guo Liu, Yanyan Guo, Zhiqiang Wang “Multiaspect sentiment analysis
for Chinese online social reviews based on topic modeling and HowNet lexicon.”
Knowl-Based Syst 2013;37:186–95.
30. Martın-Valdivia, Marıa-Teresa, Martınez-Camara Eugenio, Perea-Ortega Jose-M,
Alfonso Urena-Lopez L. “Sentiment polarity detection in Spanish reviews combining
supervised and unsupervised approaches” Expert Syst Appl 2013.
31. Xu, H., Zhang, F. and Wang, W. (2015) “Implicit feature identification in Chinese
reviews using explicit topic mining model” Knowledge-Based Systems, 76, pp.166-175.
32. Poria, S., Cambria, E., Ku, L.W., Gui, C. and Gelbukh, A. (2014) “A rule-based
approach to aspect extraction from product reviews” In Proceedings of the second
workshop on natural language processing for social media (SocialNLP) (pp. 28-37).
33. Chatterji, S., Varshney, N. and Rahul, R.K. (2017) “AspectFrameNet: a frameNet
extension for analysis of sentiments around product aspects” The Journal of
Supercomputing, 73(3), pp.961-972.