Multilingual Sentiment Analysis in Malayalam
Multilingual Sentiment Analysis in Malayalam
by
April,2024
i
MULTILINGUAL SENTIMENT ANALYSIS
by
April, 2024
ii
iii
School of Computer Science and Engineering
CERTIFICATE
Date:
Date:
(Seal of SCOPE)
iv
ABSTRACT
v
vi
TABLE OF CONTENT
CHAPTER 1
INTRODUCTION …………………………………………. 1
CHAPTER 2
LITERATURE SURVEY (1-15) …………………………. 3
CHAPTER 3
DATASET DESCRIPTION ……………………………… 16
CHAPTER 4
PROPOSED WORK
4.1 Proposed Work Introduction ……………………….. 18
4.2 Research Objectives …………………………………. 20
4.3 Flowchart Diagram …………………………………. 22
4.4 Flowchart Explanation …………………………….... 23
4.5 Modules Explanation ……………………………….. 24
4.6 Advantages and Disadvantages of models used …… 40
CHAPTER 5
RESULT AND DISCUSSION
5.1 Evaluation Matrix ………………………………… 43
vii
5.2 Observation and Discussion
5.2.1 Tables ……………………………………… 45
5.2.2 Bar Graphs ……………………………….. 47
5.3.3 Confusion Matrices ………………………. 49
CHAPTER 6
CONCLUSION ………………………………………… 52
Appencices
CODE …………………………………………………… 53
REFERNCES …………………………………………… 58
viii
LIST OF FIGURES
1. Flowchart Diagram 22
2. YouTube Comment Originality Detection
Accuracy 45
3. YouTube Comment Originality Detection 45
Accuracy on English translated Dataset
4. Fake News Detection Accuracy 46
5. Fake News Detection Accuracy on English
translated Dataset 46
6. Bar Graph for YouTube Comment Originality
Detection Accuracy 47
7. Bar Grapgh for Fake News Detection Accuracy 48
8. F1 Score for all the Models 49
ix
Chapter 1
INTRODUCTION
A significant area of research in natural language processing (NLP) is multilingual
sentiment analysis, which provides deep insights into the rich and varied fabric of
human emotion exhibited in many languages around the world. Sentiment analysis's
fundamental goal is to understand and classify the emotional undertone of written
material, a process that gets more difficult and subtle when applied to several languages.
The interpretation of sentiment can be greatly influenced by the distinctive linguistic
characteristics, cultural settings, and colloquial idioms that are peculiar to each
language, contributing to the complexity of language use.
1
This study attempts to shed light on the difficulties and effectiveness of multilingual
sentiment analysis in processing Malayalam literature by contrasting conventional
machine learning models with cutting-edge deep learning techniques. The results not
only add to the wider conversation on sentiment analysis in a variety of languages, but
they also provide useful information about how to best utilise natural language
processing (NLP) methods for Malayalam, a language that has great cultural and social
significance in the digital era.
2
Chapter 2
Literature Survey
2.1Deep Learning Model for COVID-19 Sentiment Analysis on Twitter
In the landscape of sentiment analysis research, the work by Salvador Contreras
Hernández, María Patricia Tzili Cruz, and José Martín Espínola stands out for its timely
focus on public sentiment during the COVID-19 pandemic as expressed in tweets
originating from Mexico. The significance of their study lies in the objective to
understand the public's reaction to the pandemic, which has had unprecedented global
impact. By targeting social media, specifically Twitter, the authors acknowledge the
platform's role as a contemporary barometer for public opinion.
The methodology of the study is noteworthy for its application of BERT-based models
tailored to the Spanish language, employing a semi-supervised learning framework.
This choice is significant as it delves into the linguistic nuances inherent in Spanish
tweets, which may be inadequately captured by models trained on data from other
languages or contexts. By benchmarking these specialized models against both
multilingual BERT variants and traditional machine learning classifiers, the research
provides a robust comparative analysis, highlighting the capabilities and limitations of
each approach within the scope of sentiment analysis.
The findings of this research are particularly compelling; Spanish language models
demonstrated superior precision in sentiment analysis over their multilingual and
traditional counterparts. This precision is not just a technical victory but carries
profound implications for public health strategies, suggesting that fine-tuned language
models can yield insights into public sentiment with high accuracy. These insights have
the potential to inform public health policy and decision-making, making sentiment
analysis a valuable tool in the ongoing management of the COVID-19 pandemic and in
preparing for future public health crises. The study sets a precedent for the importance
of language-specific models in sentiment analysis and their practical application in
shaping responsive and informed public health interventions.
3
2.2 Multilingual text categorization and sentiment analysis
In the collaborative work of George Manias, Argyro Mavrogiorgou, and Athanasios
Kiourtis, the complexities of classifying Twitter's multilingual content are examined.
The study's aim is to advance solutions that transcend specific domains and language
barriers, an endeavor crucial in today’s interconnected digital landscape.
The findings of the study offer a dual perspective: multilingual BERT-based classifiers
are proficient, but zero-shot classification stands out for its scalable application across
languages, despite occasionally lagging in accuracy behind fine-tuned models. This
insight provides a valuable direction for future research in multilingual natural language
processing, balancing between broad applicability and precision.
4
methodology allowed for a dynamic representation of the evolution of public interest
and sentiment, providing an illustrative timeline of the pandemic's progression as seen
through the eyes of Facebook users worldwide.
The conclusion of their study is particularly insightful, offering a unique window into
the ebb and flow of public opinion and the key subjects of discussion during the early
and intense months of the COVID-19 crisis. By capturing the chronological
development of the pandemic's public discourse in a multilingual context, the study not
only enriches the understanding of social media dynamics but also serves as a vital tool
for policymakers and health officials. It underscores the importance of leveraging social
media analytics to gauge public sentiment, which can guide response strategies and
communication plans during global health emergencies.
The methodology of the study involved a comprehensive assessment of both static and
Transformer-based embeddings, utilizing various datasets and classifiers tailored to the
Portuguese language. The goal was to meticulously identify which type of word
representation could yield the most accurate results in sentiment analysis within the
context of Brazilian Portuguese social media interactions.
The results of the research are quite revealing, indicating that the BERTimbau model,
which was specifically trained to understand the nuances of the Portuguese language,
significantly surpassed other models. This bespoke BERT model's superior
performance in sentiment classification underscores the value of language-specific
training in natural language processing tasks. The findings serve as a testament to the
potential of fine-tuned language models in enhancing the precision of sentiment
5
analysis, especially for languages other than English, which often lack the depth of
research and resources.
The study's conclusion is resoundingly positive, indicating that their proposed model
significantly outstrips current leading methods in the field of fake news detection. This
advancement is a testament to the capsule neural network’s potential for transforming
the landscape of multilingual text analysis, demonstrating a powerful tool for more
accurately and efficiently identifying false information across the world's languages.
This research not only marks a step forward in combating fake news but also illustrates
the growing necessity for advanced computational models that can navigate the
complexities of language on a global scale.
6
Youness Madani, Mohammed Erritali, and Belaid Bouikhalene have developed a
novel sentiment analysis methodology tailored to evaluate Moroccan tweets concerning
COVID-19, integrating a recommender system approach to dissect and categorize
public sentiment. Their innovative technique addresses the critical need for
understanding public opinion in times of crisis, especially within specific national
contexts.
The study's methodology employs collaborative filtering, leveraging four key tweet
features alongside sentiment analysis tools—specifically, the SenticNet dictionary and
the TextBlob Python library. This approach signifies an adaptation of recommender
system strategies, commonly seen in e-commerce, to the domain of sentiment analysis,
aiming to identify and track the collective emotional pulse of a nation through its social
media output.
Hate speech online is not just about the presence of certain keywords or phrases; it's
about context, cultural nuance, and language. The authors’ approach with HateDetector
acknowledges this complexity by incorporating an improved seagull optimization
7
algorithm to extract features from the data, which is crucial for understanding the
subtleties of language that distinguish hate speech from benign communication.
Alongside this, a hybrid diagonal-gated recurrent neural network is employed for both
detection and sentiment analysis, allowing for a nuanced interpretation of text data. The
combination of these two advanced methodologies indicates a deep understanding of
both the technical and social dimensions of hate speech.
The study's results are significant: HateDetector has shown considerable improvements
in accuracy, precision, recall, and F-measure over existing models. These metrics are
vital in the field of natural language processing and sentiment analysis, and
improvements in these areas suggest that HateDetector is capable of identifying hate
speech more reliably than ever before. In essence, the tool does not just find hateful
content; it understands it, and this understanding is crucial for any meaningful
moderation strategy on social platforms. It’s a promising development for social media
platforms that grapple with the tide of multilingual hate speech, providing a more
reliable way to monitor and manage the online dialogue to foster safer digital
environments.
The paper serves as a compendium, delving into a variety of deep learning models such
as Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and
Generative Adversarial Networks (GANs), which have each been adapted to address
different dimensions of sentiment analysis. The authors examine the specific utilities of
these models, evaluating their performance in intricate tasks that include, but are not
limited to, aspect-based sentiment analysis, which focuses on extracting sentiment with
8
respect to specific aspects of a service or product; multilingual sentiment analysis,
which grapples with sentiment expressed across different languages; and multimodal
sentiment analysis, which interprets sentiment from combined textual, audio, and visual
cues. This multifaceted review provides insights into the strengths and limitations of
each model, offering a clear view of the current state of the art.
In conclusion, the authors affirm the potent capabilities of deep learning in sentiment
analysis but also recognize the necessity for ongoing innovation. They emphasize the
importance of developing more advanced, task-specific deep learning methods to
enhance the accuracy and effectiveness of sentiment analysis. By addressing existing
challenges and outlining potential avenues for future research, the paper not only
celebrates the achievements of deep learning in this area but also serves as a clarion call
for the continuation of inventive research efforts to refine these methods further. This
is pivotal for the evolution of sentiment analysis, ensuring it remains adept at handling
the increasing complexity and nuances of human emotion expressed across diverse and
ever-expanding digital platforms.
9
the original texts and their translations to ensure a thorough analysis. Feature
visualization was accomplished using SHAP (SHapley Additive exPlanations),
enhancing interpretability of the data. They propose an innovative approach that
amalgamates three top-performing classifiers—Vader, Amazon BERT, and Sent140
BERT—aiming to leverage the strengths of each to produce a more robust sentiment
analysis tool.
The conclusion drawn from this research acknowledges the valuable insights sentiment
analysis can offer in media studies, yet it does not shy away from discussing the
inherent challenges associated with multilingual and domain-specific datasets. By
fusing multiple classifiers and incorporating explainable AI techniques, the authors
suggest that the accuracy and reliability of sentiment analysis can be significantly
improved. One of the most intriguing outcomes of this study is the identification of a
dichotomy in the representation of Olympic legacies, oscillating between utopian and
dystopian narratives. This dichotomy not only reflects the multifaceted impact of the
Olympics on host cities but also highlights the power of media in shaping public
perception and discourse. Through this work, Mello, Cheema, and Thakkar contribute
a nuanced understanding of how complex events are represented across cultural and
linguistic divides, and how advanced computational techniques can better capture the
subtleties of global narratives.
10
Leveraging data from the DravidianLangTech@ACL2022 shared task, the research
team employed a variety of models including logistic regression, CNN, Bi-LSTM, and
the aforementioned transformer models, assessing them against metrics of accuracy,
precision, recall, and F1-score. These models were augmented with sophisticated word
embedding techniques, like GloVe and BERT, to capture the semantic richness of code-
mixed language more effectively. The study presents a comprehensive comparative
analysis, culminating in the finding that the adapter-BERT model, tailored for the
specific context of code-mixed data, notably outshines others in terms of performance.
It achieves 65% accuracy in sentiment analysis and 79% in identifying offensive
language, setting a new benchmark for such tasks.
The results underscore the effectiveness of customizing pre-trained models to suit the
unique aspects of code-mixed data, a task that has become increasingly relevant as
linguistic data becomes more diverse and complex. The researchers advocate for future
exploration into multitask learning paradigms and enhancements to dataset quality,
aiming to refine the tools for handling the intricacies of mixed-language text. This work
not only propels forward the field of sentiment analysis in linguistically diverse settings
but also acts as a catalyst for future innovations that could further improve the precision
and applicability of natural language processing tools.
11
extracted features are then processed through a novel neural network architecture they
introduced—a hybrid diagonal gated recurrent neural network (H-DGRNN)—which is
tailored for the dual tasks of hate speech detection and sentiment analysis.
The study's conclusive evidence demonstrates notable strides in model performance,
showing significant gains in accuracy, precision, recall, and the F-measure. These
metrics indicate not only the effectiveness of the proposed method but also its potential
to serve as a powerful tool in the current landscape of social media platforms, where
multilingual communication and hate speech are pervasive. By enhancing the accuracy
of detecting hate speech and analyzing sentiment in code-mixed text, the study by Kar
and Debbarma provides a critical contribution to creating safer online environments
and advancing the field of natural language processing for multilingual content.
12
presence of the Hindi language online. They outline potential future directions for
research, underscoring the importance of this work in understanding and harnessing the
vast wealth of sentiment expressed by Hindi speakers on digital platforms. This review
not only provides a roadmap for future explorations but also serves as a benchmark for
the current state of sentiment analysis in Hindi, paving the way for innovations that
could significantly impact both academic research and practical applications in the
realm of natural language processing.
Spanning an impressive volume of over 2.2 billion tweets, the dataset encompasses a
range of languages, thereby offering a panoramic view of global sentiment and
discussions. This broad linguistic range allows for an inclusive understanding of the
pandemic's social impact worldwide. The data is further enriched with sentiment
analysis and named entity recognition algorithms, equipping researchers with the tools
needed to perform a nuanced analysis of the conversations and to identify key themes,
stakeholders, and sentiments within the vast swathes of Twitter data.
The conclusion posited by Lopez and Gallemore is that this curated dataset stands as a
critical resource for academics and professionals striving to grasp the evolution of
public sentiment throughout the pandemic. Its value lies not just in its size but in its
potential to serve as a foundational platform for a variety of social media data analyses,
from tracking the spread of information to understanding the public’s emotional and
cognitive responses to the pandemic's progression. The dataset offers a unique
opportunity to chronicle the pandemic's narrative as told through the eyes of millions
on social media, presenting an unfiltered account of humanity's collective experience
and reaction to COVID-19.
13
2.14 Persian Text Sentiment Analysis Based on BERT and Neural
Networks
Siroos Rahmani Zardak, Amir Hossein Rasekh, and Mohammad Sadegh Bashkari
have conducted a pivotal study in the field of natural language processing, focusing
specifically on sentiment analysis within the Persian language. Their work underscores
a significant gap in computational linguistics—effective sentiment analysis for Persian
text—and seeks to bridge this by adapting the Bidirectional Encoder Representations
from Transformers (BERT) algorithm for Persian language sentiment analysis.
The conclusion of the study confirms the superiority of the BERT algorithm in
analyzing sentiment in Persian texts. The BERT model achieves high accuracy and F1
scores, outstripping earlier methods. This significant improvement not only
demonstrates BERT's adaptability across languages but also indicates a promising
direction for future research in sentiment analysis for underrepresented languages.
The methods deployed in this research are comprehensive, employing both machine
learning and deep learning techniques. These include Support Vector Machines (SVM),
14
Naive Bayes (NB), AdaBoost, Multi-Layer Perceptrons (MLP), Logistic Regression
(LR), Random Forest (RF), one-dimensional Convolutional Neural Networks (CNN-
1D), Long Short-Term Memory networks (LSTM), Bidirectional LSTMs (Bi-LSTM),
Gated Recurrent Units (GRU), Bidirectional GRUs (Bi-GRU), and a fine-tuned version
of the multilingual BERT (mBERT) model.
In their conclusion, the researchers highlight the superior performance of the mBERT
model, which utilizes BERT's pre-trained word embeddings. The mBERT model
achieved an impressive F1 score of 81.49%, showcasing its proficiency in interpreting
and analyzing sentiment in the Urdu language.
15
Chapter 3
Dataset-Description
This study utilizes two distinct datasets sourced from the CodaLab Fake News
Detection in Dravidian Languages competition, specifically designed for the Dravidian-
LangTech@EACL 2024. The datasets are instrumental in exploring sentiment analysis
and fake news detection within the Malayalam language, representing a significant step
towards understanding and analyzing Dravidian languages through advanced natural
language processing techniques.
The first dataset comprises 3,200 Malayalam comments, derived from various
YouTube videos, intended to distinguish between original and fake content. This
dataset is split into a 70:30 ratio for training and validation purposes, resulting in 2,240
entries for training and the remainder for validation. An additional set of 1,020
comments is reserved for testing. Each entry consists of two fields:
• Text: The comment text in Malayalam, presenting a diverse array of topics and
sentiments.
The second dataset focuses on Malayalam news articles, comprising 1,700 entries
divided similarly into a 70:30 training-validation split, yielding 1,190 entries for
training and the rest for validation. The testing set includes 250 news articles. The
dataset features two main columns:
16
2.3 Helsinki-NLP Pipeline and English Translation
To enhance the analysis and applicability of NLP tools traditionally optimized for
English, the study incorporates the Helsinki-NLP pipeline for machine translation. This
approach involves converting the Malayalam datasets into English, thereby enabling
the use of a broader range of NLP models and tools that may not be directly applicable
to Malayalam due to language-specific constraints.
17
Chapter 4
Proposed Work
The capacity to precisely evaluate sentiment and identify bogus news in multilingual
environments has become critical in the rapidly changing world of digital
communication. There is an urgent need for sophisticated natural language processing
(NLP) systems that can handle the complexities of language-specific nuances, cultural
contexts, and a wide range of sentiment expressions due to the proliferation of
information and the increasing significance of regional languages on digital platforms.
This study suggests a thorough approach to deal with these issues in the millions of
speakers of the Dravidian language Malayalam, which is rich in linguistic variety.
The suggested method investigates sentiment analysis and false news detection in
Malayalam by utilising both conventional machine learning algorithms and state-of-
the-art deep learning approaches. This study takes a multipronged strategy to determine
the best techniques for identifying and classifying opinions and the accuracy of news
reports in Malayalam. In addition to improving the state of NLP research on Dravidian
languages, the system is meant to offer useful tools and insights that may be used in a
variety of contexts, such as social media monitoring and content moderation, among
others.
Our method is based on analysing Malayalam texts for sentiment and authenticity using
a variety of machine learning classifiers, such as Support Vector Machine (SVM),
Random Forest (RF), Logistic Regression, and Naive Bayes. The system investigates
the effectiveness of deep learning models concurrently, concentrating on the BERT
classifier and its modifications for the Malayalam language via direct application,
translation-based methods, and sophisticated tokenization strategies. Through the use
of both classic and new NLP approaches, this dual-path exploration seeks to improve
accuracy and provide a deeper understanding of the sentiment and authenticity of
Malayalam texts.
18
The suggested approach would evaluate these models' performance in relation to
important indicators using a strict evaluation framework, providing a thorough grasp of
their advantages and disadvantages. This methodical approach not only lays the
foundation for applying these findings to other languages and circumstances, but it also
promises to improve our understanding of sentiment analysis and fake news detection
in Malayalam.
The suggested method is proof of the potential of fusing various NLP techniques as we
traverse the challenges of multilingual sentiment analysis and fake news identification.
By providing a strong framework for evaluating sentiment and authenticity in
Malayalam, it seeks to make a substantial contribution to the subject and, in the process,
further knowledge of digital communication in multilingual societies.
19
4.2 Research Objectives
The main objective of this work is to improve the field of natural language processing
(NLP) by creating and assessing models that can identify fake news and perform
sentiment analysis in Malayalam, a language that is widely spoken in the digital sphere
but is not well-represented in NLP studies. The following are the study's particular
goals:
Examine how well the classic machine learning classifiers—Random Forest (RF),
Logistic Regression, Support Vector Machine (SVM), and Naive Bayes—identify
feelings and identify false news in Malayalam text. In order to achieve this goal,
these models will be compared in order to determine which method performs best
while considering performance indicators like accuracy, precision, recall, and F1
score.
Examine the use of deep learning models for Malayalam sentiment analysis and
fake news identification, with a focus on the BERT classifier. This covers the use
of translated texts through the Helsinki-NLP pipeline, direct application on
Malayalam texts, and enhanced Malayalam text tokenization using the Indic-BERT
tokenizer.
20
4.2.4 To Assess the Impact of Language Translation on Model
Performance
Identify the main challenges and limitations encountered when applying NLP
models to Malayalam text. This objective aims to provide insights into language-
specific issues, data scarcity, and the nuances of cultural context in sentiment
analysis and fake news detection.
21
4.3 Flowchart Diagram
Fig 1
22
4.4 Flowchart Explanation
4.4.1 Dataset
This is the primary repository of raw data collected for analysis. In this context,
it consists of Malayalam comments and news articles that will be the subject of
sentiment analysis and fake news categorization. The dataset serves as the
foundational input for all subsequent steps in the flowchart.
Recognizing the limitations in NLP tools for Malayalam, this stage involves
translating the dataset into English. The translation aims to leverage the
extensive range of tools and models developed for English, potentially
increasing the accuracy of sentiment analysis. This step is a testament to the
multilingual aspect of the project and addresses the challenge of language
disparity in NLP resources.
Post-translation, the dataset exists in a form that is more accessible to NLP tools
predominantly designed for the English language. This translated corpus retains
the original semantics but is now suitable for processing by models trained on
English datasets.
23
4.4.4 Preprocessing
• Tokenization: This is the practice of dividing text into tokens, which can
be words, phrases, symbols, or other meaningful elements. Tokenization
is essential for models to analyse and understand the text at a granular
level.
24
4.4.6 Model Selection
4.4.7 Evaluation
25
4.5 Modules Explanation
4.5.1 Vectorization
• Count Vector-ization
26
• Word Embeddings
Once the data is vectorized, it is ready for use in machine learning models for
various tasks, including classification, clustering, and sentiment analysis.
4.5.2 Tokenization
27
• Implementing Tokenization: Use a tokenization method. For
many languages, simple space-based tokenization may be
sufficient, but for languages like Malayalam with agglutinative
properties, a more sophisticated approach may be required.
28
• Subtokenization: For models like BERT, subtokenization is
often used where words are broken down into smaller pieces that
allow the model to handle a wider variety of words it hasn't seen
before.
29
4.5.3 Translation (Malayalam to English)
Translation in the context of NLP involves converting text from one language
to another, in this case, from Malayalam to English. It's a complex task that
requires understanding and preserving the meaning of the source text. Here's
how the translation process unfolds in detail:
30
• Quality Assurance: Implement quality assurance measures, such as
back-translation (translating the text back to the original language) or
human review, especially for critical applications.
Often used to solve classification problems, the Support Vector Machine (SVM)
is a strong and adaptable supervised machine learning technique that may be
applied to regression as well as classification applications. It is especially
helpful in situations when precise decision boundaries are required because of
its capacity to identify the hyperplane that most effectively partitions a dataset
into classes. Because of its versatility in handling different kinds of data and its
efficiency in high-dimensional environments, this approach is particularly well-
liked.
31
Using Support Vector Machines (SVM) to analyse datasets for plagiarism
detection can be quite helpful in determining the similarities and differences
between texts or documents. Word frequency, n-grams, and semantic similarity
are just a few of the aspects that may be analysed using SVM to determine if a
document is authentic or plagiarised. The SVM-generated decision planes
facilitate the differentiation of authentic content from instances of plagiarism,
which in turn allows for the detection and mitigation of content reuse or
unauthorised copying.
32
Steps to Perform SVM:
33
outperform more sophisticated algorithms and are especially useful for text
classification tasks such as spam detection or sentiment analysis.
34
4.5.6 Logistic Regression
Logistic Regression is also a mathematical and stats related method for dividing
a more than one dataset independent variables that determine an result. The
result event is measured with a dich-otomous variable. It is used extensively in
fields like medicine and social sciences, as well as machine learning, where it
serves as a baseline algorithm for binary classification problems.
35
3. Model Fitting: Fit the logistic regression model to the dataset,
estimating the probability of the dependent variable given the
independent variables.
4. Model Validation: Validate the model's performance with metrics
such as ROC curve, AUC, or confusion matrix.
5. Optimization: Use techniques like regularization to prevent
overfitting and to enhance the model's generalization capabilities.
6. Prediction: Employ the model to estimate the probability of the
target class for new data points.
36
6. Ensemble Prediction: Make predictions by averaging the output of
all the individual trees (regression) or by majority voting
(classification).
7. Error Reduction: Achieve error reduction through the aggregation of
less correlated trees.
37
4.5.8 Bi-directional Encoder Representation from Transformer
(BERT)
38
2: Preprocess the Text Data
3: Tokenization
4: Input Representation
BERT requires three kinds of input data: input IDs (token IDs from the
tokenizer), attention masks (which tokens should be attended to), and
token type IDs (used to distinguish different sentences).
5: Fine-tuning BERT
Pass your input data through BERT and add a classification layer on top.
Train the model on your labeled dataset. During training,
backpropagation is used to fine-tune the pre-trained BERT parameters
and learn the weights of the added output layer.
6: Evaluation
After fine-tuning, evaluate the model on a validation set to check its
performance. Use relevant metrics like accuracy, precision, recall, F1-
score, etc.
39
7: Inference
Use the fine-tuned BERT model to make predictions on new data. The
[CLS] token's output embedding can be used as the aggregate sequence
representation for classification tasks.
Advantages:
• Speed: It's known for its speed in training and prediction, making it
highly efficient for large datasets.
• Simplicity: Its algorithmic simplicity and the assumption of feature
independence make it straightforward to implement and understand.
• Performance: Despite its simplicity, Naive Bayes can perform
remarkably well in text classification tasks such as spam detection and
document categorization.
Disadvantages:
40
4.6.2 Random Forest (RF)
Advantages:
Disadvantages:
• Complexity and Size: Models can become large and unwieldy, requiring
significant memory storage.
• Computation Time: Training time can be long, especially as the number
of trees increases.
Advantages:
Disadvantages:
41
4.6.4 Support Vector Machine (SVM)
Advantages:
Disadvantages:
• Scalability: Training time can grow quickly with the size of the data,
making it less suitable for large datasets.
• Parameter Tuning: Requires careful tuning of parameters, including the
choice of kernel and regularization parameters, which can be complex
and time-consuming.
Advantages:
Disadvantages:
42
Chapter 5
5.1.1 Accuracy
One of the easiest metrics to employ when assessing
categorization models is accuracy. It calculates the percentage of
all predictions that the model accurately predicts, including true
positives and true negatives:
43
against the model's predictions, providing insights into the types
of errors made by the model.
The matrix is especially useful for binary classification tasks
and can be extended to multiclass classification problems.
A confusion matrix for binary classification has four parts:
By using bar graphs, you can quickly identify which models are
performing well and which ones are lagging in certain areas. For
instance, you might plot the accuracy of several models on a bar
graph to see which model has the highest accuracy. You can also
use bar graphs to visualize the performance of a single model
across different metrics, making it easier to assess the model's
strengths and weaknesses comprehensively.
44
5.2 Observation and Explanation
Below are the tables summarizing the performance (Accuracy and Training Time) of
each model for the YouTube Comment Originality dataset and the Fake News
Detection dataset.
• Training Time is the time it took to train each model. BERT-based models took
significantly longer due to their complex architectures, with
EnglishTranslation+BERT taking the longest at nearly 4 hours.
45
Fake News Detection
Figure 5 - Tabe for Fake News Detection Accuracy on English translated Dataset
Similar performance metrics were used here.
• Accuracy for this dataset varied, with Random Forest performing the best at
68.6% and BERT performing the worst at 60.52%.
• Training Time also showed a similar trend to the YouTube comments, with
BERT-based models requiring more time, indicating the complexity and
computational demands of these models.
46
Bar Graph for the accuracy is shown below.
47
Fig 7- Fake News Detection
48
Youtube Comment Dataset Models Confusion Matrices Result-:
Figure 9 – Table for F1 Score for Yobute Comment Originality Detection on English
Translated Dataset
49
Fake News Dataset Models Confusion Matrices Result-:
Figure 10 – Table for F1 Score for Fake News Detection on Malayalma Language
Dataset
Figure 11 – Table for F1 Score for Fake News Detection on English Translated Dataset
50
We evaluated the performance of various machine learning models on two distinct
datasets: one composed of Malayalam YouTube comments aimed at distinguishing
between original and fake content, and the other focusing on Malayalam news articles
for fake news detection. Initially, we observed that across both datasets, models such
as Naive Bayes, Random Forest, Logistic Regression, BERT, and IndicBERT exhibited
competitive accuracies. However, there was a notable discrepancy when comparing the
performance on the original Malayalam dataset versus the English translated dataset.
Interestingly, upon applying the models to the English translated dataset, we found a
notable increase in accuracy for certain models, notably Naive Bayes and Random
Forest, in both the YouTube comment originality and fake news detection tasks.
Additionally, the training time for the translated dataset was generally reduced
compared to the original Malayalam dataset. This observation suggests that while the
translation process introduces linguistic variations, it may also enhance the model's
ability to discern patterns in the data. Moreover, the reduced training time for the
translated dataset implies that the models may require less computational resources to
achieve comparable performance, potentially enhancing their scalability and efficiency
in real-world applications. These findings underscore the importance of considering
language translation techniques in multilingual text analysis tasks, offering insights into
optimizing model performance and resource utilization.
51
Chapter 6
A pivotal factor in the observed performance discrepancies is the size of the datasets.
Deep learning models, like BERT, are inherently data-hungry and designed to excel
with vast quantities of information. With 3,200 comments for YouTube and 1,700
articles for news, the datasets may be considered undersized for the deep learning
models to fully capture and generalize the nuances of the Malayalam language.
Meanwhile, the traditional models showed notable efficiency, possibly due to their
suitability for smaller datasets and less complex feature spaces.
This research thus not only sheds light on the capabilities of current sentiment analysis
models for Malayalam texts but also underscores the critical need for larger datasets to
fully exploit the prowess of deep learning techniques. Moving forward, expanding the
datasets or employing data augmentation strategies might be necessary to elevate the
performance of deep learning models and to establish a more definitive benchmarking
within the field of multilingual sentiment analysis. The study serves as a stepping stone
for future research, emphasizing the balance between model choice, dataset size, and
computational efficiency in the pursuit of accurate sentiment analysis.
52
Appendices
Appendix 1
Code
Code Snippets for all the steps involved and models used -:
Vectorization
Naïve Bayes
53
SVM
Random Forest
54
Logistic Regression
BERT Tokenizer
55
INDIC BERT Tokenizer
BERT MODEL
56
Malayalam to English Translation Coede
57
REFERNCES
[1] Contreras Hernández, S., Tzili Cruz, M. P., Espínola Sánchez, J. M., & Pérez Tzili,
A. (2023). Deep learning model for covid-19 sentiment analysis on twitter. New
Generation Computing, 41(2), 189-212.
[2] Manias, G., Mavrogiorgou, A., Kiourtis, A., Symvoulidis, C., & Kyriazis, D.
(2023). Multilingual text categorization and sentiment analysis: a comparative analysis
of the utilization of multilingual approaches for classifying twitter data. Neural
Computing and Applications, 35(29), 21415-21431.
[3] Amara, A., Hadj Taieb, M. A., & Ben Aouicha, M. (2021). Multilingual topic
modeling for tracking COVID-19 trends based on Facebook data analysis. Applied
Intelligence, 51, 3052-3073.
[4] Vianna, D., Carneiro, F., Carvalho, J., Plastino, A., & Paes, A. (2023). Sentiment
analysis in Portuguese tweets: an evaluation of diverse word representation models.
Language Resources and Evaluation, 1-50.
[5] Mohawesh, R., Maqsood, S., & Althebyan, Q. (2023). Multilingual deep learning
framework for fake news detection using capsule neural network. Journal of Intelligent
Information Systems, 60(3), 655-671.
[6] Madani, Y., Erritali, M., & Bouikhalene, B. (2023). A new sentiment analysis
method to detect and Analyse sentiments of Covid-19 moroccan tweets using a
recommender approach. Multimedia Tools and Applications, 82(18), 27819-27838.
[7] Anjum, & Katarya, R. (2023). HateDetector: Multilingual technique for the analysis
and detection of online hate speech in social networks. Multimedia Tools and
Applications, 1-28.
[8] Habimana, O., Li, Y., Li, R., Gu, X., & Yu, G. (2020). Sentiment analysis using
deep learning approaches: an overview. Science China Information Sciences, 63, 1-36.
[9] Mello, C., Cheema, G. S., & Thakkar, G. (2023). Combining sentiment analysis
classifiers to explore multilingual news articles covering London 2012 and Rio 2016
Olympics. International Journal of Digital Humanities, 5(2), 131-157.
58
[10] Shanmugavadivel, K., Sathishkumar, V. E., Raja, S., Lingaiah, T. B.,
Neelakandan, S., & Subramanian, M. (2022). Deep learning based sentiment analysis
and offensive language identification on multilingual code-mixed data. Scientific
Reports, 12(1), 21557.
[11] Kar, P., & Debbarma, S. (2023). Multilingual hate speech detection sentimental
analysis on social media platforms using optimal feature extraction and hybrid diagonal
gated recurrent neural network. The Journal of Supercomputing, 79(17), 19515-19546.
[12] Sidhu, S., Khurana, S. S., Kumar, M., Singh, P., & Bamber, S. S. (2023). Sentiment
analysis of Hindi language text: a critical review. Multimedia Tools and Applications,
1-30.
[13] Lopez, C. E., & Gallemore, C. (2021). An augmented multilingual Twitter dataset
for studying the COVID-19 infodemic. Social Network Analysis and Mining, 11(1),
102.
[14] Zardak, S. R., Rasekh, A. H., & Bashkari, M. S. (2023). Persian Text Sentiment
Analysis Based on BERT and Neural Networks. Iranian Journal of Science and
Technology, Transactions of Electrical Engineering, 47(4), 1623-1634.
[15] Khan, L., Amjad, A., Ashraf, N., & Chang, H. T. (2022). Multi-class sentiment
analysis of urdu text using multilingual BERT. Scientific Reports, 12(1), 5436.
[17] Kakwani, D., Kunchukuttan, A., Golla, S., Gokul, N. C., Bhattacharyya, A.,
Khapra, M. M., & Kumar, P. (2020, November). IndicNLPSuite: Monolingual corpora,
evaluation benchmarks and pre-trained multilingual language models for Indian
languages. In Findings of the Association for Computational Linguistics: EMNLP 2020
(pp. 4948-4961).
59
BERT contributes to advancements in sentiment analysis and fake news detection by providing a robust framework that captures context and nuances in multiple languages. Its ability to fine-tune for specific languages or contexts enhances model precision. For Malayalam, a BERT classifier was applied both on original Malayalam text and English-translated content, allowing for comparative insights into language-specific performance . Moreover, the BERT-based models' superior performance in Spanish sentiment analysis during the COVID-19 pandemic illustrates its effectiveness in language-specific scenarios . BERT's architecture, which allows bidirectional context understanding, enables precise sentiment analysis and detection tasks across different linguistic data .
Cultural and linguistic contexts significantly affect the effectiveness of NLP models for languages like Malayalam and Hindi. The unique linguistic features and cultural nuances embedded within these languages can impact how sentiment is expressed and interpreted. For Malayalam, translation into English for model application can lead to loss of cultural context, necessitating a balance between using advanced English-focused NLP tools and maintaining linguistic authenticity . In the case of Hindi, the ongoing refinement of resources like Hindi SentiWordNet indicates the challenges in capturing sentiment accurately with existing tools, accentuated by the language’s complex syntactic and semantic structures .
The effectiveness of BERT classifier configurations for Malayalam text varies based on the approach taken. Models using an Indic-BERT tokenizer demonstrate improved handling of the Malayalam script, potentially enhancing tokenization fidelity and context understanding. Conversely, BERT classifiers on English-translated text may leverage superior English-focused NLP tools but can suffer from translation-related context loss. Direct application of BERT on original Malayalam text retains linguistic and cultural authenticity but may face limitations due to less mature tool support . The comparative use of these configurations allows for balanced insights into model performance across linguistic boundaries while highlighting inherent trade-offs in accuracy and cultural representation .
Multilingual social media data analysis has significant potential for informing public health strategies, as evidenced by its role during the COVID-19 pandemic. Studies utilizing platforms like Twitter and Facebook illustrate how public sentiment and discourse can be effectively analyzed across multiple languages. For example, Spanish-specific sentiment analysis during COVID-19 yielded immediate insights with high precision that are critical for public health policy adjustments . Similarly, multilingual topic modeling on Facebook data aided in tracking pandemic trends and discourse evolution across seven languages, providing a dynamic understanding essential for responsive health interventions . These analyses underscore the strategic advantage of leveraging social media analytics in global health crises for real-time sentiment detection and public communication refinement.
Deep learning models like BERT offer enhanced understanding of context and language nuances compared to traditional machine learning approaches such as SVM and Naive Bayes when applied to sentiment analysis and fake news detection for Malayalam. BERT's ability to process linguistic intricacies through bidirectional reading allows for a deeper semantic comprehension, leading to more accurate classification outcomes . In contrast, traditional methods rely on simpler feature-based selection and are potentially limited by vectorization and tokenization step constraints . While logistical and computational complexity may be higher for deep learning models, their adaptability and superior performance justify these challenges in complex linguistic contexts like Malayalam .
The studies indicate several challenges in multilingual sentiment analysis, such as the scarcity of NLP tools for less represented languages like Malayalam and Hindi, which affects the accuracy and reliability of sentiment analysis. Specifically for Malayalam, translation to English using the Helsinki-NLP pipeline revealed trade-offs between linguistic authenticity and model accessibility . In Hindi, the evolution of resources like Hindi SentiWordNet suggests ongoing efforts to refine sentiment analysis techniques, highlighting the challenges in addressing linguistic nuances and ensuring the effectiveness of sentiment analysis models in these languages .
NLP techniques face several challenges in languages with fewer digital resources, such as Malayalam and Hindi. The scarcity of annotated datasets, limited lexical resources, and often less mature NLP tool support hinder the development of effective sentiment analysis models for these languages. This results in a dependence on translation to more resource-rich languages like English, which risks losing cultural and contextual subtleties . These challenges underscore the need for specialized model architectures and data preprocessing methods tailored to the linguistic structures of underrepresented languages. The implications are significant: addressing these gaps can lead to more inclusive digital interaction and support the preservation of linguistic diversity in global digital spaces .
Strategies to enhance NLP tasks for code-mixed multilingual data include customizing pre-trained models and incorporating multitask learning paradigms. One approach involves the use of hybrid neural networks, such as the hybrid diagonal gated recurrent neural network, which effectively handles the intricacies of mixed-language text . These strategies improve sentiment analysis accuracy and help address challenges in detecting subtle linguistic features, such as offensive language, in code-mixed data. By doing so, they advance the applicability and precision of NLP tools in diverse and linguistically complex environments, impacting areas like social media moderation and sentiment analysis .
The findings suggest that despite zero-shot classification's broad applicability across multiple languages, it sometimes falls short in accuracy compared to fine-tuned BERT-based models. This implies a trade-off between the ease of scalability in zero-shot models and the precision of language-specific fine-tuned models. The insight offers a balanced approach for future research in multilingual natural language processing: leveraging zero-shot models for quick deployment across languages while employing fine-tuning for critical precision-centric tasks like sentiment analysis .
The study by Contreras Hernández et al. emphasizes the significance of language-specific models in sentiment analysis by showcasing that Spanish language models outperformed both multilingual and traditional classifiers. This precision was crucial for accurately capturing public sentiment regarding the COVID-19 pandemic, which has significant implications for informing public health strategies. By using BERT-based models tailored to the Spanish language, the study illustrates that fine-tuned language models can provide insights into public sentiment with high accuracy, setting a precedent for the necessity of language-specific models in effectively analyzing sentiment across different linguistic contexts .