Analysis of Bollywood
Boycott Trend
PROJECT GUIDE: SUBMITTED BY:
Dr. G.L. Prajapati Akshat Garg
Deepanshu Panwar
Prakrit Jain
Abdul samad
2
INTRODUCTION
In India one of the most trending topic is boycotting of
bollywood movies because of nepotism, suicide cases of
various talented celebrities and above all, subpar movies.
There is a war going on between bollywood and general
public regarding the issue. Among the General public, there
are also people who support boycott, oppose boycott and
people who have maintained a neutral stance.
3
Reason for choosing the Project
Bollywood is the name given to the Indian film industry focused
on Hindi Language movies and based in Mumbai. Recently,
due to various reasons like Subpar movie scripts, controversial
topics and nepotism, the bollywood is facing a severe backlash
in the form of a Bollywood Boycott trend which is taking social
media platforms like twitter with fire. These trends have a
definite effect on earnings and profits of recent bollywood
movies and that is exactly what we want to find out.
4
Objective
The objective of this project is to extract tweets from the social
media platform “Twitter” and analyze them to find out the
stance of general public on the recent movies and what effects
does these tweets and resulting sentiments of public has on the
reputation and earnings of these movies.
5
Techstack & Tools used
▸ PYTHON : Language used for implementation
▸ Scikit Learn : Python Library used for preprocessing and implementing Machine
Learning Models.
▸ VADER : Used for Sentiment Intensity Analysis of Tweets
▸ WORDCLOUD : library for Constructing Word Cloud of the tweets
▸ NLTK LIBRARY : stopwords removal , Tokenisation and Maximum Entropy
6
Steps Involved
▸ Following Are the Steps involved in the Analysis:
7
Data Cleaning and Preprocessing
We followed the below steps to get our data cleaned and ready
for our analysis :
1 – Remove all the links and urls from the text.
2 – Remove the tagged user (‘@…’)
3 – Remove the Hashtags(‘#…’) from the text and store them in a
seperate column. Theymight be useful.
8
Data Cleaning and Preprocessing
4 – Remove all the punctuations from the text. They generally
don’t contribute to the sentiment of a text.
5 – Lower the case of letters in the text.
6 – Remove the stopwords from the text.
7 – This is the last step which is tokenizing the cleaned text. This
step will also help in removing tweets which are in languages
other than English.
9
Data Cleaning and Preprocessing
This Pre-Processing and cleaning process, results in a reduction of
the number of tweets, the number of tweets is decresased
drastically,
Number of 310079
Tweets Before
Cleaning
Number of 145883
Tweets After
Cleaning
10
Polarity Score Using VADER
VADER sentimental analysis relies on a dictionary that maps
lexical features to emotion intensities known as sentiment
scores. The sentiment score of a text can be obtained by
summing up the intensity of each word in the text. We used it
calculate how positive, negative or neutral a tweet is.
11
Polarity Score Using VADER
The application of VADER lead ot the following coompositon of
tweet categories.
CATEGORY PERCENT (COUNT)
Positive 38.1% (65445)
Negative 44.9% (55553)
Neutral 17.1% (24885)
12
Exploratory Data Analysis
Exploratory Data Analysis refers to the critical process of
performing initial investigations on data so as to discover
patterns,to spot anomalies,to test hypothesis and to check
assumptions with the help of summary statistics and graphical
representations.
13
Exploratory Data Analysis
BAG OF WORDS
We will start our exploaration by creating a Bag of Words. The bag-of-words model is
a way of representing text data when modeling text with machine learning
algorithms. The following word cloud is obtained,
14
Exploratory Data Analysis
HAHSTAG ANALYSIS
The hashtags on a site like twitter is a way for people to state the topic about which
their tweet is about. The following is a graph showing the top ten hashtags.
15
Exploratory Data Analysis
ASSIGNING CATEGORIES TO TWEETS
• Now,let us use our polarity score to categorize these tweets in three categories.
Those who ‘Support Boycott’, those who ‘Oppose Boycott’ and those who are
neutral. For this we manually observed many tweets carefully and came to a
conclusion that we can assign the categories on the basis of polarity score only.
After categorizing the tweet, the following data is obtained.
CATEGORY PERCENT (COUNT)
Oppose Boycott 38.1% (65445)
Support Boycott 44.9% (55553)
Neutral 17.1% (24885)
16
Exploratory Data Analysis
ASSIGNING CATEGORIES TO TWEETS
• The above categories can be better visualized as a pie chart,
17
Machine Learning
Now that we have our dataset ready, we can move towards
applying machine learning. But before applying machine
learning, in order to train our models we need to do two
basic things,
1. Encode our target
2. Vectorize the cleaned tweets.
18
Machine Learning
1. Encoding our Target:
Why do we need to Encode the target? Well, to answer that we need
to understand that machines don’t understand human language. It
understand machine languages and numbers, therefore we need
Encoding. We will use label Encoder for the same purpose. The
following encoding is done for the data,
CATEGORY ENCODING
Oppose Boycott +1
Support Boycott -1
Neutral 0
19
Machine Learning
1. Vectorization:
Why we need to vectorize the cleaned tweets? Vecorization will transform the text
into
into meaningful
meaningful representation
representation of of integers
integers or or numbers
numbers which
which is used
is used to to
fit fit
machine
machine
learning learningfor
algorithm algorithm for predictions.
predictions.
For this purpose we are going to use TF-IDF vectorizer. TF-IDF Vectorizer is a
measure
measure of of originality
originality of of a word
a word by by comparing
comparing thethe number
number of of times
times a word
a word appears
appears in with
in document document with the
the number number of documents
of documents the word
the word appears in. appears in. The
The following
following
formula formula
is used is used to
to calculate thecalculate
value of the value
a word of atf-idf.
using word using tf-idf.
20
Machine Learning
NAIVE BAYES :
Naive Bayes classifiers are a collection of classification algorithms based on
Bayes’ Theorem. It is not a single algorithm but a family of algorithms where
all of them share a common principle, i.e. every pair of features being
classified is independent of each other. The naive bayes algorithm is based on
the following formula,
where, A is class variable and B is a dependent feature vector with dimension
d i.e. B = (b1,b2,b2,…, bd), where d is the number of
variables/features of the sample.
21
Machine Learning
NAIVE BAYES :
The Flow Diagram of Naive Bias Algorithm is as follows,
22
Machine Learning
MAXIMUM ENTROPY ALGORITHM :
The Max Entropy classifier is a probabilistic classifier which belongs
to the class of exponential models. Unlike the Naive Bayes at we
discussed in the previously, the Max Entropy does not assume that
the features are conditionally independent of each other. The MaxEnt
is based on the Principle of Maximum Entropy and from all the
models that fit our training data, selects the one which has the
largest entropy. The Max Entropy classifier can be used to solve a
large variety of text classification problems such as language
detection, topic classification, sentiment analysis and more.
23
Machine Learning
MAXIMUM ENTROPY ALGORITHM :
The Flow diagram of MaxEnt Is as Follows,
24
Model Evaluation
To measure the performance of our model we need some evaluation metrics on the basis
of which we are going to compare our models. For this project we will use, Accuracy,
Precision, Recall, F1-score.
Accuracy is the percentage of correctly predicted data, the equation show how to
calculate it.
Accuracy = (Tp + Tn)/(Tp+Tn+Fp+Fn)
Recall is the percentage of correctly predicted positive data, the equation show how to
calculate it.
Recall = (Tp )/(Tp+Fn)
25
Model Evaluation
Precision is the percentage of positive data predicted as positive, the equation show how
to calculate it.
Precision = (Tp + Tn)/(Tp+Fp)
F-Score is a representation of recall and precision, the equation show how to calculate it.
F1-score = (2*P*R)/(P+R)
Here,
Tp – True Positive Tn – True Negative
Fp – False Positive Fn – False Negative
P – Precision R – Recall
26
Model Evaluation
The accuracy of the models came out to be around 72.8% and 70.1% for Naive
Bayes and Maximum Entropy respectively.
MODEL ACCURACY
Naive Bayes 0.728
Maximum Entropy 0.701
27
Model Evaluation
The other metrics are calculated class wise and are depicted in the following
tables.
NAIVE BAYES :
Class Precision Recall F1 score
Support 0.65 0.69 0.67
Boycott
Oppose 0.71 0.6 0.65
Boycott
Neutral 0.74 0.7 0.72
28
Model Evaluation
The other metrics are calculated class wise and are depicted in the following
tables.
MAXIMUM ENTROPY :
Class Precision Recall F1 score
Support 0.69 0.7 0.69
Boycott
Oppose 0.74 0.66 0.7
Boycott
Neutral 0.65 0.73 0.69
29
Conclusion
As for our Analysis on the bollywood boycott trend, we observed that majority
of the population wants to boycott the Bollywood. To reinforce our findings
lets take examples of two recent major bollywood movies Laal Singh Chaddha
and Vikram Vedha. Laal Singh Chaddha was tagged as poor remake of
hollywood movie forest gump while Vikram Vedha was also not liked much by
the population as compared to its ol counterpart. The following table depicts
the ratings given by two major reviw Website IMDB and Rotten Tomatoes(RT),
MOVIE IMDB RT VERDICT
Laal Singh Chaddha 5.5/10 0.65 Flop
Vikram Vedha 7.1/10 0.79 Below Average
30
Conclusion
This can be inferred from above data that the Boycott trend had significant
effect on the success of the movie, both Popularity wise and monetry wise,
in a negative way.
31
Conclusion
Now, let’s talk about our machine learning model. We chose Naive Bayes
and Maximum Entropy classifier for our machine learning part and found out
that Naive Bayes and Maximum Entropy both showed a similar performance.
Either Naïve Bayes and Maximum Entropy as classification method could be
used for textual data.
If we looked from number alone, Naïve Bayes has an advantage for about
1.7% better than Maximum Entropy, but with Maximum Entropy has an
attribute of non-independent, we could disregard the little difference and have
more reliable result generated by Maximum Entropyclassification model.
THANK YOU