0% found this document useful (0 votes)
6 views47 pages

Telugu Offensive Language Detection Report

The project report details the development of an offensive language identification system tailored for Telugu, addressing the challenges posed by code-mixed and transliterated text. Utilizing advanced machine learning and natural language processing techniques, the project enhances classification accuracy through data augmentation and deep learning models. The outcome is a web application that flags harmful comments in real-time, contributing to safer online environments for Telugu-speaking communities.

Uploaded by

kannanvishruth
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views47 pages

Telugu Offensive Language Detection Report

The project report details the development of an offensive language identification system tailored for Telugu, addressing the challenges posed by code-mixed and transliterated text. Utilizing advanced machine learning and natural language processing techniques, the project enhances classification accuracy through data augmentation and deep learning models. The outcome is a web application that flags harmful comments in real-time, contributing to safer online environments for Telugu-speaking communities.

Uploaded by

kannanvishruth
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

A

Project Report
On
Offensive Language Identification in Telugu
Submitted for partial fulfillment of the requirements for the award of the degree
Of
BACHELOR OF ENGINEERING
In
COMPUTER SCIENCE AND ENGINEERING
By
Mr. Pradyumna Chacham(2451-21-733-066)
Mr. Koushik Pyarasani (2451-21-733-083)
Mr. Deekshith Sripadhi (2451-21-733-086)
Under the guidance of
Mr. B. Venkataramana
Assistant Professor
Department of CSE

MATURI VENKATA SUBBA RAO(MVSR) ENGINEERING COLLEGE


Department of Computer Science and Engineering
(Affiliated to Osmania University & Recognized by AICTE)
Nadergul, Saroor Nagar Mandal, Hyderabad – 501510
Academic Year: 2024-25
Maturi Venkata Subba Rao Engineering College
(Affiliated to Osmania University, Hyderabad)
Nadergul(V), Hyderabad-501510

CERTIFICATE
This is to certify that the project work entitled “Offensive Language Identification in
Telugu” is a bonafide work carried out by Mr. Chacham Pradyumna (2451-21-733-
066), Mr. Koushik Pyarasani (2451-21-733-083) and Mr. Deekshith Sripadhi (2451-
21-733-086)in partial fulfilment of the requirements for the award of degree of
Bachelor of Engineering in Computer Science and Engineering from Maturi
Venkata Subba Rao (MVSR) Engineering College, affiliated to OSMANIA
UNIVERSITY, Hyderabad, during the Academic Year 2023-24 under our guidance and
supervision.

The results embodied in this report have not been submitted to any other university or
institute for the award of any degree or diploma to the best of our knowledge and belief.

Internal Guide Head of Department


Mr. B Venkataramana Prof. J Prasanna Kumar
Assistant Professor Professor
Department of CSE Department of CSE
MVSREC. MVSREC.
External Examiner

i
DECLARATION
This is to certify that the work reported in the present project entitled “Offensive
Language Identification in Telugu” is a record of bonafide work done by us in the
Department of Computer Science and Engineering, Maturi Venkata Subba Rao
(MVSR) Engineering College, Osmania University during the Academic Year 2024-
25. The reports are based on the project work done entirely by us and not copied from
any other source. The results embodied in this project report have not been submitted
to any other University or Institute for the award of any degree or diploma.

Mr. Pradyumna Chacham Mr. Koushik Pyarasani Mr. Deekshith Sripadhi


2451-21-733-066 2451-21-733-083 2451-21-733-086

ii
ACKNOWLEDGEMENT
We would like to express our sincere gratitude and indebtedness to our project guide
Mr. B. Venkataramana for his valuable suggestions and interest throughout the course
of this project.

We are also thankful to our principal Dr. Vijaya Gunturu, and Mr. J Prasanna
Kumar, Professor and Head, Department of Computer Science and Engineering,
Maturi Venkata Subba Rao Engineering College, Hyderabad for providing excellent
infrastructure for completing this project successfully as a part of our B.E. Degree
(CSE). We would like to thank our project coordinator for her constant monitoring,
guidance and support.

We convey our heartfelt thanks to the lab staff for allowing us to use the required
equipment whenever needed. We sincerely acknowledge and thank all those who gave
directly or indirectly their support in the completion of this work.

Mr. Pradyumna Chacham (2451-21-733-066)


Mr. Koushik Pyarasani (2451-21-733-083)
Mr. Deekshith Sripadhi (2451-21-733-086)

iii
VISION
• To impart technical education of the highest standards, producing competent
and confident engineers with an ability to use computer science knowledge to
solve societal problems.
MISSION
• To make learning process exciting, stimulating and interesting.
• To impart adequate fundamental knowledge and soft skills to students.
• To expose students to advanced computer technologies in order to excel in
engineering practices by bringing out the creativity in students.
• To develop economically feasible and socially acceptable software.
PEOs:
PEO-1: Achieve recognition through demonstration of technical competence for
successful execution of software projects to meet customer business objectives.
PEO-2: Practice life-long learning by pursuing professional certifications, higher
education or research in the emerging areas of information processing and intelligent
systems at a global level.
PEO-3: Contribute to society by understanding the impact of computing using a
multidisciplinary and ethical approach.
PROGRAM OUTCOMES (POs)
At the end of the program the students (Engineering Graduates) will be able to:
1. Engineering knowledge: Apply the knowledge of mathematics, science,
engineering fundamentals, and an engineering specialisation for the solution of
complex engineering problems.
2. Problem analysis: Identify, formulate, research literature, and analyse complex
engineering problems reaching substantiated conclusions using first principles
of mathematics, natural sciences, and engineering sciences.
3. Design/development of solutions: Design solutions for complex engineering
problems and design system components or processes that meet the specified
needs with appropriate consideration for public health and safety, and cultural,
societal, and environmental considerations.
4. Conduct investigations of complex problems: Use research-based knowledge
research methods including design of experiments, analysis an interpretation of
data, and synthesis of the information to provide valid conclusions.

iv
5. Modern tool usage: Create, select, and apply appropriate techniques,
resources, and modern engineering and IT tools including prediction and
modelling to complex engineering activities with an understanding of the
limitations.
6. The engineer and society: Apply reasoning informed by the contextual
knowledge to assess societal, health, safety, legal, and cultural issues and the
consequent responsibilities relevant to the professional engineering practice.
7. Environment and sustainability: Understand the impact of the professional
engineering solutions in societal and environmental contexts, and demonstrate
the knowledge of, and the need for sustainable development.
8. Ethics: Apply ethical principles and commit to professional ethics and
responsibilities and norms of the engineering practice.
9. Individual and teamwork: Function effectively as an individual, and as a
member or leader in diverse teams, and in multidisciplinary settings.
10. Communication: Communicate effectively on complex engineering activities
with the engineering community and with the society at large, such as being
able to comprehend and write effective reports and design documentation, make
effective presentations, and give and receive clear instructions.
11. Project management and finance: Demonstrate knowledge and understanding
of the engineering and management principles and apply these to one’s work,
as a member and leader in a team, to manage projects and in multidisciplinary
environments.
12. Lifelong learning: Recognise the need for and have the preparation and ability
to engage in independent and life-long learning in the broadest context of
technological change.
PROGRAM SPECIFIC OUTCOMES (PSOs)
13. (PSO-1) Demonstrate competence to build effective solutions for computational
real-world problems using software and hardware across multi-disciplinary
domains.
14. (PSO-2) Adapt to current computing trends for meeting the industrial and
societal needs through a holistic professional development leading to pioneering
careers or entrepreneurship.

v
COURSE OBJECTIVES AND OUTCOMES
Course Code: U21PW881CS
Course Objectives
• To enhance practical and professional skills.
• To familiarize tools and techniques of systematic Literature survey and
documentation.
• To expose the students to industry practices and team work.
• To encourage students to work with innovative and entrepreneurial ideas.
Course Outcomes
Upon completion of the course, the student will be able to:
1. Demonstrate the ability to synthesize and apply the knowledge and skills
acquired in the academic program to real-world problems.
2. Evaluate different solutions based on economic and technical feasibility.
3. Effectively plan a project and confidently perform all the aspects of project
management.
4. Demonstrate effective written and oral communication skills.
5. Present the proposed project using PPT.

vi
ABSTRACT
The increasing prevalence of offensive language on social media platforms has
raised significant concerns about online safety and inclusivity. Detecting offensive
content in low-resource languages like Telugu poses a unique challenge due to limited
annotated datasets and linguistic tools. This project focuses on the detection of
offensive language in Telugu using the HOLD dataset, which includes code-mixed,
transliterated, and pure Telugu text. Initially, the dataset comprised 4,000 training and
500 testing instances. After refining the dataset to include only Telugu text, the dataset
size was reduced to 1,336 training and 242 testing samples, setting the stage for more
advanced model development.

To improve classification accuracy, we first applied data augmentation techniques such


as backtranslation, synonym replacement, and paraphrasing, which significantly
increased the diversity of the dataset, expanding it to 12,904 samples. We then
leveraged feature extraction methods like Word2Vec and FastText to capture semantic
relationships and contextual information within the text. Finally, deep learning models,
including transformer-based architectures like BERT and its variants, were employed
to capture the complexities of code-mixed and transliterated data, improving model
performance and handling nuances in Telugu text.

The outcome of this project is an automated offensive language detection system,


integrated into a web application that identifies and flags harmful comments in real-
time. By combining advanced data preprocessing, augmentation techniques, and state-
of-the-art deep learning models, this project not only advances natural language
processing for Telugu but also contributes to fostering safer and more respectful digital
environments. The solution's scalability and adaptability can be extended to other low-
resource languages, offering a significant impact on moderating online content and
promoting inclusivity.

vii
TABLE OF CONTENTS
Certificate i
Declaration ii
Acknowledgment iii
Vision & Mission iv
PEOs,Pos and PSOs iv-v
Course Objectives and Outcomes vi
Abstract vii
Table of Contents viii
List of Figures ix
List of Tables ix

CONTENTS
1. INTRODUCTION 1
1.1 Problem Statement 1
1.2 Objective 2
1.3 Motivation 2
1.4 Scope of the Project 3
2. LITERATURE SURVEY 6
3. SYSTEM DESIGN 8
3.1 Flowchart 8
3.2 System Architecture 8
3.3 ML Module 10
3.3.1 Machine Learning Models 11
3.3.2 Deep Learning Models 12
3.3.2 Transformer Models 12
3.4 Voting Mechanism and Predictions 13
4.1 Dataset Processing and Augmentation 14
4.2 Implementation of Each Module 19
[Link] AND RESULT 23
[Link] AND FUTURE ENHANCEMENTS 30
REFERENCES 31
APPENDIX (SOURCE CODE) 32

viii
LIST OF FIGURES
Figure 3. 1 Data Flow 8
Figure3 2 System Architecture 9

LIST OF TABLES
Table 4.1 Dataset Statistics 17
Table 5.3 Testing Results on Augmented Dataset 25
Table 5.4 Comparison of Results for Original and Augmented Datasets 28

ix
Offensive Language Identification in Telugu

1. INTRODUCTION
The proliferation of offensive language on social media platforms has become
a pressing concern, impacting both individual users and societal well-being. Detecting
such content in low-resource languages like Telugu is particularly challenging due to
limited datasets and linguistic tools. Telugu, widely spoken in India, faces significant
gaps in NLP research, especially with the prevalence of code-mixed text, which
complicates the identification of harmful language. This project aims to address these
challenges by developing an offensive language detection system tailored to Telugu.

The project leverages advanced machine learning models and NLP techniques,
including text preprocessing, feature extraction with Word2Vec and FastText, and deep
learning models based on transformer architectures. These models help capture the
linguistic nuances of Telugu and code-mixed content, enhancing the accuracy of
offensive language classification. The success of this project has the potential to impact
various stakeholders, including researchers, developers, and policymakers, by
providing an effective and scalable solution tool for managing offensive language and
improving online safety on digital platforms.

1.1 Problem Statement


The rise of social media has revolutionized communication, offering unprecedented
opportunities for self-expression and interaction. However, it has also led to an increase
in the spread of offensive language, creating challenges for maintaining respectful
online spaces. This issue is particularly prominent in low-resource languages like
Telugu, where the absence of large annotated datasets, specialized linguistic tools, and
tailored models hinders the development of effective offensive language detection
systems.

In addition, the use of code-mixed text—where Telugu is interspersed with


English or other languages—adds another layer of complexity. Traditional natural
language processing (NLP) techniques struggle to handle such mixed-language content,
making the identification of offensive language even more difficult. These challenges
make it essential to develop a system capable of accurately detecting offensive language

Dept. of CSE, MVSREC 1


Offensive Language Identification in Telugu

in Telugu, addressing the unique linguistic features of the language and its code-mixed
variations.

1.2 Objective
The primary objective of this project is to design and implement a system for
detecting offensive language in Telugu text, focusing on code-mixed, transliterated, and
pure Telugu content. This involves preprocessing and refining the HOLD dataset to
accurately handle these variations of Telugu text, ensuring it reflects real-world
language usage. The project aims to enhance classification accuracy by applying
advanced data augmentation techniques, such as backtranslation, synonym substitution,
and paraphrasing. Additionally, a web application will be developed to flag offensive
user comments in real-time, providing immediate feedback to users. Ultimately, the
project seeks to contribute to the broader field of natural language processing by
addressing the unique challenges of processing low-resource languages like Telugu,
advancing the ability to detect offensive content in such contexts.

1.3 Motivation
The increasing reliance on online communication for both personal and
professional interactions has made the detection of offensive language more critical
than ever. The harmful effects of toxic language extend beyond disrupting meaningful
conversations, also having serious psychological implications for users. Despite the
growing need, much of the research and development in offensive language detection
has been tailored to resource-rich languages like English, leaving low-resource
languages such as Telugu underserved. This gap in available tools and datasets for
Telugu makes it harder to address harmful content in one of the most widely spoken
languages in India. By focusing on the challenges posed by Telugu, particularly its
code-mixed and transliterated forms, this project aspires to bridge this gap and create a
safer online space for Telugu-speaking communities. Moreover, it offers an opportunity
to explore advanced techniques in natural language processing and machine learning,
fostering both technical innovation and practical problem-solving in the field of low-
resource language processing.

Dept. of CSE, MVSREC 2


Offensive Language Identification in Telugu

1.4 Scope of the Project


This project focuses on addressing the challenge of detecting offensive language
in Telugu text, specifically in its code-mixed, transliterated, and pure forms. Given the
widespread use of Telugu on social media platforms, this project aims to develop a
solution that can accurately classify offensive content in various textual formats
commonly encountered online. The scope includes preprocessing the HOLD dataset to
handle these different forms of Telugu, followed by applying advanced data
augmentation techniques to enhance the dataset and improve model accuracy. The
project also involves the development of a web application that provides real-time
detection of offensive language, offering immediate feedback to users. While the
primary focus is on Telugu, the methods and techniques explored in this project have
the potential to be extended to other low-resource languages with similar linguistic
challenges. The project also contributes to the field of natural language processing by
addressing the gaps in existing tools for underrepresented languages, thereby enhancing
the ability to detect harmful content in diverse linguistic contexts.

1.5 Software and Hardware Requirements

Software Requirements:

1. Python: Python is the core programming language for the backend of the
project, providing robust support for data processing, machine learning, and
natural language processing (NLP). Python is used to implement the machine
learning models for offensive language detection and integrate them into the
web application.

2. Flask: Flask is used as the web framework for the backend. It is a lightweight
and scalable framework that facilitates the development of web APIs for real-
time offensive language detection. Flask serves as the interface for handling
requests between the frontend (React) and the machine learning models,
ensuring smooth communication and model deployment.

3. React: React is used for the frontend development of the web application. It
allows for the creation of a dynamic and responsive user interface where users

Dept. of CSE, MVSREC 3


Offensive Language Identification in Telugu

can input text and receive immediate feedback on whether their comments are
offensive. React enables the creation of reusable components and ensures an
interactive and smooth user experience.

4. Scikit-learn: Scikit-learn is used for implementing traditional machine learning


classifiers such as Support Vector Machines (SVMs), Logistic Regression, and
Random Forest. It is essential for preprocessing, training, and evaluating
machine learning models, providing an efficient framework for
experimentation.

5. TensorFlow or PyTorch: TensorFlow and PyTorch are utilized for deep


learning model development. These frameworks support the training and fine-
tuning of complex neural network architectures, including transformers like
BERT, to capture the nuances of Telugu and code-mixed text in offensive
language detection.

6. Pandas and NumPy: Pandas is used for data manipulation, including loading,
cleaning, and structuring datasets. NumPy accelerates numerical operations,
making it essential for matrix computations and data transformations in machine
learning pipelines.

7. Natural Language Toolkit (NLTK): NLTK is essential for text preprocessing,


which includes tokenization, stemming, lemmatization, and stopword removal.
These steps ensure that the input text is standardized and ready for machine
learning models.

8. Jupyter Notebook or Integrated Development Environment (IDE): Jupyter


Notebook is employed for prototyping, experimentation, and visualization of
machine learning models. For managing the project codebase, IDEs such as
PyCharm or Visual Studio Code are used to provide a robust development
environment for both backend and frontend components.

9. [Link]: [Link] is required to run the React application, enabling the


development of the frontend web interface. It allows for the execution of

Dept. of CSE, MVSREC 4


Offensive Language Identification in Telugu

JavaScript on the server side and is necessary for the development environment
of the React app.

10. Web Browser: A modern web browser (e.g., Google Chrome, Mozilla Firefox)
is required to test and interact with the deployed web application. The frontend
and backend will be tested and accessed via a browser, providing a smooth
interface for end-users.

Hardware Requirements:

1. Processor: A multi-core processor, preferably Intel Core i5 or higher, is


necessary to handle the computational tasks associated with data preprocessing,
model training, and running both the frontend and backend.

2. RAM: At least 8 GB of RAM is recommended to run the project efficiently,


especially during the model training and web application deployment phases.
For optimal performance, 16 GB or more is preferable.

3. Storage: A minimum of 20 GB of free disk space is needed for storing datasets,


dependencies, and model files. An SSD is preferred to ensure faster read/write
speeds, which improves performance during model training and data
manipulation.

4. GPU: A dedicated GPU (e.g., NVIDIA GTX 1060 or higher) is recommended


for accelerating deep learning model training. A GPU significantly speeds up
the training process for large datasets, especially when using neural network
architectures.

Dept. of CSE, MVSREC 5


Offensive Language Identification in Telugu

2. LITERATURE SURVEY

[Link] Year of Author(s) Technique Summary Limitation


Publication

1 2023 Abhinaba Logistic Explored logistic Limited


Bala et al Regression regression exploration of
with combined with transformer-
embeddings embeddings for based models
abusive comment for improved
detection in accuracy in
code-mixed Telugu
Tamil and language
Telugu. settings.

2 2021 Nikhil Pretrained Applied Did not focus


Kumar et transformers transformer- on Telugu, and
al (mBERT, based models for experiments
XLM-R) offensive lacked code-
language mixed text
detection in variations.
Tamil,
Malayalam, and
Kannada.

3 2024 Nafisa Transliteration Proposed Dependent on


Tabassum augmentation transliteration- transliteration
et al with Telugu- based data quality, limiting
BERT augmentation for applicability to
code-mixed noisy or highly
Telugu text, mixed datasets.
achieving

Dept. of CSE, MVSREC 6


Offensive Language Identification in Telugu

superior
accuracy.

4 2020 Premjith B Character N- Used character Struggles with


et al grams with N-grams as input semantic
CNN features for CNN nuances in
to detect offensive
offensive content language,
in Dravidian particularly in
languages. Telugu.

5 2024 Namit Fine-tuned Demonstrated the Lacked


Khanduja mBERT and effectiveness of experiments on
et al XLM-R multilingual code-mixed
transformers for Telugu datasets,
monolingual limiting real-
Telugu hate world
speech detection. applicability.

Dept. of CSE, MVSREC 7


Offensive Language Identification in Telugu

3. SYSTEM DESIGN
3.1 Flowchart
The flowchart outlines the step-by-step process of the Offensive Language Detection
System for Telugu text, including code-mixed and transliterated inputs. The process
starts with User Input, where the user submits a comment. The text is then processed in
the Preprocessing phase, removing unwanted noise such as special characters, emojis,
and URLs, while handling code-mixed and pure Telugu text through transliteration and
tokenization. Next, Feature Extraction converts the cleaned text into numerical
representations using techniques like TF-IDF, Word2Vec, and FastText, capturing its
semantic and syntactic properties. These features are analyzed by a combination of
machine learning and deep learning models, including SVM, Naive Bayes, and
transformer-based architectures like mBERT and XLM-R. Finally, the system performs
language classification, categorizing the text as offensive or non-offensive and
providing real-time feedback to the user. This workflow ensures efficient and accurate
detection of offensive content in Telugu, especially in code-mixed and transliterated
forms.

Figure 3. 1 Data Flow

3.2 System Architecture


The Offensive Language Detection System architecture is designed to process and
classify user-generated Telugu text, including code-mixed and transliterated data. The
system employs a combination of machine learning (ML), deep learning, and
transformer-based models to ensure accurate detection of offensive and non-offensive

Dept. of CSE, MVSREC 8


Offensive Language Identification in Telugu

language. The architecture leverages state-of-the-art feature extraction techniques, such


as Word2Vec, FastText, and TF-IDF, enabling efficient text representation and
classification. A modular design ensures the system can scale and be deployed in real-
time for content moderation on various platforms.

Figure3 2 System Architecture

Input Text
The Input Text refers to the raw user-generated content, which may be in pure
Telugu, code-mixed Telugu-English, or transliterated formats. The system processes
this input through various stages, beginning with preprocessing to handle the unique
linguistic challenges of Telugu text, including transliteration and tokenization.
Feature Extraction
Feature Extraction is crucial to transforming the unstructured text into
structured, machine-readable formats. The system uses Word2Vec to capture semantic

Dept. of CSE, MVSREC 9


Offensive Language Identification in Telugu

relationships between words, FastText to manage rare words and subword information,
and TF-IDF to measure the importance of terms within the dataset. These features
provide the necessary input for the classification models that follow.
TF-IDF
TF-IDF (Term Frequency-Inverse Document Frequency) is a popular technique
in text representation that measures the importance of a word in a document relative to
its frequency in a corpus. Term Frequency (TF) captures how often a word appears in
a document, while Inverse Document Frequency (IDF) reflects how rare or unique a
word is across all documents. Words that are frequent in a document but rare in the
corpus receive higher weights, helping to highlight important terms. TF-IDF is useful
in offensive language detection to identify key words associated with harmful content,
improving the accuracy of classification models.
Word2Vec
Word2Vec is a technique for generating word embeddings, which are dense vector
representations of words that capture semantic relationships based on their context in
the text. Word2Vec uses a shallow neural network to learn these word representations
from large corpora of text by predicting neighboring words (CBOW model) or
predicting a word given its context (Skip-gram model). These word embeddings are
used as input features for downstream tasks such as text classification and sentiment
analysis.
FastText
FastText is an extension of Word2Vec that improves word embeddings by
considering subword information. It represents each word as a bag of character n-
grams, which helps capture the meaning of rare or out-of-vocabulary words by
understanding their subword components. FastText is especially useful for languages
with rich morphology, like Telugu, where words can change forms based on suffixes
or prefixes.

3.3 ML Module
The ML Module in the system architecture is designed to handle the
classification of Telugu text using a combination of traditional machine learning

Dept. of CSE, MVSREC 10


Offensive Language Identification in Telugu

(ML) models, deep learning models, and transformer-based models. Each of these
models offers distinct strengths in terms of handling the complexity of code-mixed,
transliterated, and pure Telugu text.
3.3.1 Machine Learning Models
Random Forest
Random Forest is an ensemble learning algorithm that constructs a multitude of
decision trees during training and outputs the class that is the mode of the classes (for
classification problems). It works by creating multiple decision trees and combining
their results to improve the model's accuracy and prevent overfitting. Random Forest is
known for handling large datasets and maintaining accuracy even when some of the
data is missing.
Logistic Regression
Logistic Regression is a statistical model used for binary classification problems. It
uses the logistic function to model the probability of a certain class (offensive or non-
offensive) based on input features. This model is simple, interpretable, and effective for
problems where the relationship between the features and the output is linear.
Support Vector Machine
Support Vector Machine (SVM) is a supervised machine learning algorithm
effective for classification tasks, especially with high-dimensional data. It works by
finding the hyperplane that best separates classes while maximizing the margin between
them. This is valuable for offensive language detection, where features like word
frequencies can be high-dimensional.
Decision Tree
A Decision Tree is a flowchart-like structure in which each internal node
represents a "test" or "decision" on an attribute (e.g., whether a word in the text is
offensive), and each branch represents the outcome of that test. The leaf nodes represent
the final classification. It is an intuitive model for classification, where the goal is to
segment the input space into regions of similar outcomes. However, decision trees are

Dept. of CSE, MVSREC 11


Offensive Language Identification in Telugu

prone to overfitting, which is why ensemble methods like Random Forest are used to
improve accuracy.
3.3.2 Deep Learning Models
Convolutional Neural Network (CNN)
CNNs are a class of deep neural networks commonly used in image processing and
sequential data analysis. In NLP, CNNs apply filters (kernels) over the input text to
capture local features, such as phrases or word patterns, which can help in detecting
specific linguistic cues. This makes CNNs effective for text classification tasks,
especially when detecting specific patterns of words or phrases in the text.
Bidirectional Long Short-Term Memory Networks (BiLSTM)
BiLSTM is a type of recurrent neural network (RNN) designed to capture long-
range dependencies in sequential data. Unlike traditional RNNs, BiLSTM processes
data in both forward and backward directions, allowing it to retain information from
both past and future contexts within a sequence. This makes BiLSTM particularly
useful for tasks such as sentiment analysis, where the meaning of a word can depend
on both previous and subsequent words in the text.
3.3.2 Transformer Models
BERT (Bidirectional Encoder Representations from Transformers)
BERT is a transformer-based model designed to pre-train deep bidirectional
representations by jointly conditioning on both left and right context in all layers. BERT
is pre-trained on a large corpus of text and fine-tuned on task-specific datasets, which
allows it to achieve state-of-the-art performance in NLP tasks such as text classification,
question answering, and named entity recognition. BERT's bidirectional nature allows
it to capture richer context than unidirectional models, making it highly effective in
understanding the nuances of language.
mBERT (Multilingual BERT)
mBERT is a multilingual version of BERT, pre-trained on text from multiple
languages. It allows the model to handle text in different languages without requiring
separate models for each language. mBERT is useful in multilingual environments like
Telugu, as it can learn from data in both the native language and code-mixed contexts

Dept. of CSE, MVSREC 12


Offensive Language Identification in Telugu

with other languages (e.g., English), making it highly adaptable for tasks in low-
resource languages.
IndicBERT
IndicBERT is a language model specifically designed for Indian languages,
including Telugu. It is based on BERT and fine-tuned to understand the unique
linguistic characteristics of languages from the Indian subcontinent. IndicBERT is
trained on a large corpus of multilingual text in various Indian languages and is highly
effective at handling the diverse scripts and code-mixed text commonly seen in social
media data.
SBERT (Sentence-BERT)
SBERT is a modification of the BERT architecture designed to generate fixed-size
sentence embeddings that can be used for tasks like text similarity and clustering. By
fine-tuning BERT with a Siamese network structure, SBERT improves the performance
of tasks where understanding sentence-level relationships is essential. SBERT is
particularly effective for semantic textual similarity tasks and can capture nuanced
differences between offensive and non-offensive content.
XLM-R (Cross-lingual RoBERTa)
XLM-R is a transformer model designed for cross-lingual understanding,
capable of handling text in multiple languages. It is pre-trained on 100 languages and
excels in tasks that require understanding text in different languages. While it performs
well on multilingual datasets, its performance may not be as specialized as models like
IndicBERT for individual Indian languages like Telugu. However, its ability to handle
multiple languages makes it useful for cross-lingual applications.

3.4 Voting Mechanism and Predictions


To enhance classification accuracy, the system uses a Voting Mechanism. Each
model in the ML Module contributes its prediction, and the final output is determined
by the majority. If most models classify the text as offensive, the result is offensive;
otherwise, it’s non-offensive. This approach leverages the strengths of different models,
improving overall performance.

Dept. of CSE, MVSREC 13


Offensive Language Identification in Telugu

4. SYSTEM IMPLEMENTATION & METHODOLOGIES


4.1 Dataset Processing and Augmentation
The dataset used in this project primarily consisted of user-generated comments
in Telugu, including both code-mixed and pure Telugu text. Given the complexity of
code-mixed and transliterated text, along with the limited size of the dataset, data
augmentation techniques were applied to improve model performance and
generalization. Pure Telugu instances were first removed from the dataset to focus on
code-mixed and transliterated forms. The filtered dataset, now containing only code-
mixed Telugu data, formed the foundation for further augmentation.

Data augmentation is a critical step, particularly in low-resource languages like


Telugu, where labeled data is often scarce. It involves generating synthetic variations
of the existing data to enhance the dataset, reduce overfitting, and improve the model’s
ability to generalize across different linguistic patterns. In this project, three key
augmentation techniques were applied: Synonym Replacement, Paraphrasing, and
Backtranslation. These methods were applied specifically to the filtered dataset, which
contained only Telugu text. The goal was to increase the variety of Telugu language
expressions while preserving their semantic meaning and label consistency.

Synonym Replacement

Synonym replacement was performed at the lexical level, where specific words in
a sentence were substituted with semantically similar counterparts. This enhanced the
diversity of the vocabulary in the dataset and allowed the model to better understand
different lexical expressions of the same meaning.

• Word Embeddings: FastText word embeddings trained on Telugu were used to


identify semantically similar words.

• Process: A random selection of words within each sentence was chosen for
replacement, with synonyms selected based on cosine similarity in the
embedding space.

• Output: This process was repeated multiple times for each sentence, generating
diverse variations.

Dept. of CSE, MVSREC 14


Offensive Language Identification in Telugu

Paraphrasing

Paraphrasing was applied to rephrase sentences while preserving their original


meaning. Unlike synonym replacement, paraphrasing modified the sentence structure,
word order, and syntax, introducing significant diversity at the sentence level.

• Model Used: The google/byt5-base transformer model was used to generate


paraphrased versions of the Telugu sentences.

• Process: Each sentence was prefixed with "paraphrase:" and input into the
model, generating syntactically restructured sentences through beam search
decoding.

• Language Detection: A language detection mechanism ensured that only Telugu


sentences were processed.

Backtranslation

Backtranslation involved translating a sentence into a target language (English) and


then re-translating it back into Telugu. This process helped generate alternative
sentence forms, maintaining the original meaning but with different lexical and
syntactic structures.

• Process: Each Telugu sentence was first translated into English using the
mtranslate library, then translated back into Telugu.

• Output: This generated diverse paraphrases that introduced natural variations in


phrasing and structure.

Transliteration

After applying the data augmentation techniques, the resulting filtered Telugu-only
dataset was transliterated into Roman script to enhance accessibility and compatibility
with NLP tools that require Latin-script input. This step also supported downstream
tasks where script normalization is beneficial.

• Tool Used: The indic-transliteration Python library was used to perform the
transliteration.

Dept. of CSE, MVSREC 15


Offensive Language Identification in Telugu

• Process: Each Telugu sentence was converted into the IAST (International
Alphabet of Sanskrit Transliteration) format.

Combining the Datasets

The final training dataset was created by combining multiple augmented


datasets. The original dataset (containing both code-mixed and pure Telugu records)
had pure Telugu records removed, leaving a code-mixed dataset. The augmented
Telugu dataset, which included the sentences with synonym replacement, paraphrasing,
and backtranslation applied, was added to the code-mixed dataset. Additionally, the
transliterated augmented Telugu dataset was created to further expand the diversity of
the data.

The final training set was constructed by combining the following:

• The code-mixed dataset and transliterated original dataset (after removing pure
Telugu records from the original dataset).

• The augmented Telugu dataset (after applying synonym replacement,


paraphrasing, and backtranslation).

• The transliteration of the augmented Telugu dataset.

Figure 4.1 Dataset Processing

Dept. of CSE, MVSREC 16


Offensive Language Identification in Telugu

The statistics of the datasets after processing and augmentation are as follows:

Table 4.1 Dataset Statistics

Dataset Type No. of Samples

Original 4,000

Filtered (Only Telugu) 1,336

Synonym Replacement 1276

Paraphrased 1336

Backtranslated 1336

Final Telugu Dataset 5284

Transliterated 5284

Final Training Dataset (after removing duplicates) 12,904

Test 499

Dept. of CSE, MVSREC 17


Offensive Language Identification in Telugu

Dataset split:

Class Name : non-hate Class Name : hate


Number of Comments:5088
Number of Comm ents:7816 Number of Words:39671
Number of Words:66047 Number of Unique Words:15861
Number of Unique Words:20595 Most Frequent Words:
Most Frequent Words:
ni 215
chala 428 ra 199
anna 386 ki 180
చాలా 386 lo 170
i 313 miru 142
ఈ 312 mi 133
miru 294 kuda 132
మీరు 290 మీరు 123
ga 120
సూపర్277
nuvvu 120
supar 260
garu 256
Total Number of Unique Words:31762

Dept. of CSE, MVSREC 18


Offensive Language Identification in Telugu

Figure 4.3 a Figure 4.3 b

Figure 4.3 a and b: Length-freqeuency distribution of [Link] set and b. testing set

Figure 4.4 Sample instances from augmented Dataset

4.2 Implementation of Each Module


The Offensive Language Detection System is implemented using several key
stages, from data preprocessing to model training and prediction. The architecture
integrates multiple models, including traditional machine learning models, deep
learning models (CNN, BiLSTM), and transformer-based models (mBERT,
IndicBERT, SBERT). Each part of the system has been designed to process text
efficiently, leveraging the strengths of different algorithms to detect offensive language
in Telugu, especially in code-mixed and transliterated formats.

Data Preprocessing

The first step involves data preprocessing, where the input text is cleaned and
prepared for further analysis. The function clean_text() removes unwanted characters,
such as special symbols, emojis, and URLs, ensuring that the text is suitable for feature

Dept. of CSE, MVSREC 19


Offensive Language Identification in Telugu

extraction. The tokenize() function splits the text into individual tokens, and stopwords
(irrelevant common words) are removed using the function remove_stopwords().
Additionally, for handling code-mixed and transliterated text, the function
handle_code_mixed() is used to normalize text that contains both Telugu and English
components.

Feature Extraction

After preprocessing, the system extracts features from the cleaned text to be
used by the models. The TF-IDF vectorizer is employed using the
tfidf_vectorizer.fit_transform() function, which creates a sparse matrix that represents
the frequency of terms in each document, weighted by their inverse document
frequency across the corpus. Additionally, Word2Vec and FastText embeddings are
used to convert words into dense vectors that capture their semantic meaning. The
function word2vec_model() trains the Word2Vec model on the text corpus, while the
function fasttext_model() does the same for FastText, focusing on rare and out-of-
vocabulary words. These embeddings are fed as input features to the machine learning
and deep learning models.

Model Training

The system uses a combination of traditional machine learning models, deep


learning models, and transformer-based models to classify the input text as either
offensive or non-offensive.

Convolutional Neural Networks (CNN) The CNN model is used to capture local
features in the text, such as specific word patterns or n-grams. The model architecture
begins with an embedding layer that converts the input tokens into dense vectors using
Word2Vec or FastText embeddings. The Conv1D(filters=256, kernel_size=3) layer
applies convolutional filters across the word embeddings, learning local patterns that
might indicate offensive content. This is followed by a MaxPooling1D(pool_size=2)
layer, which reduces the dimensionality of the output, highlighting the most important
features. The final fully connected layer (Dense(units=1, activation='sigmoid')) outputs
a binary prediction of whether the text is offensive or not. The CNN is effective for

Dept. of CSE, MVSREC 20


Offensive Language Identification in Telugu

detecting specific sequences of words, which are useful for identifying patterns in
offensive content.

Bidirectional Long Short-Term Memory Networks (BiLSTM): BiLSTM is used to


capture long-term dependencies and contextual relationships in the text by processing
it in both forward and backward directions. The Bidirectional(LSTM(units=128,
return_sequences=True)) layer ensures that the model can learn both the past and future
context of a word in the sequence. This bidirectional approach is crucial for
understanding the meaning of words in context, especially for complex languages like
Telugu. The output of the LSTM layer is passed to a dense layer (Dense(units=1,
activation='sigmoid')), which produces the final prediction. BiLSTM models are
particularly effective for sequence-based tasks, where the meaning of a word is
influenced by both previous and subsequent words in the text.

Transformer Models (mBERT, IndicBERT, SBERT): Transformer-based


models are employed to handle the complexities of code-mixed and multilingual text.
These models, including mBERT, IndicBERT, and SBERT, use self-attention
mechanisms to process the text in parallel, allowing them to capture deep contextual
relationships between words. mBERT (Multilingual BERT) is pre-trained on a
multilingual corpus and fine-tuned on the Telugu dataset to improve its performance
for this specific task. The train_mbert() function is used to fine-tune mBERT for
offensive language detection in Telugu. Similarly, IndicBERT is fine-tuned with the
function train_indicbert() to capture the specific linguistic nuances of Indian languages.
SBERT (Sentence-BERT) generates sentence-level embeddings and is useful for
classifying entire sentences based on their semantic meaning. The function train_sbert()
fine-tunes SBERT to improve its ability to capture the meaning of sentences, making it
especially effective for understanding the context of offensive language in longer
comments.

Voting Mechanism

To enhance classification accuracy, the system implements a Voting


Mechanism that aggregates the predictions from all models. The function
aggregate_predictions() collects the results from each model—CNN, BiLSTM,

Dept. of CSE, MVSREC 21


Offensive Language Identification in Telugu

mBERT, IndicBERT, Random Forest, Logistic Regression, etc. If the majority of the
models predict the text as offensive, the final result is classified as offensive; otherwise,
it is classified as non-offensive. This mechanism ensures that the system benefits from
the diverse strengths of the different models, providing a robust decision-making
process.

Model Evaluation

The models are evaluated using a variety of metrics such as precision, recall,
F1-score, and accuracy to determine their effectiveness in classifying offensive content.
The function evaluate_model() calculates these metrics for each model, providing a
comprehensive view of their performance. The Confusion Matrix is used to visualize
the classification results, highlighting true positives, false positives, true negatives, and
false negatives, which allows for identifying areas where the models need
improvement.

Deployment

Once the models are trained and evaluated, they are integrated into a web
application for real-time detection of offensive content. The function predict_text()
processes the user input, applies the trained models, and provides immediate feedback
on whether the text is offensive or non-offensive. This real-time feedback is essential
for content moderation on platforms like social media, ensuring that harmful content is
flagged promptly.

Dept. of CSE, MVSREC 22


Offensive Language Identification in Telugu

[Link] AND RESULT

In this section, we present the evaluation of various models on the HOLD


dataset before and after applying data augmentation. The objective was to assess model
performance using key classification metrics—Precision, Recall, F1-Score, and
Accuracy—for both non-hate and hate classes. The initial evaluation was carried out
on the original HOLD dataset to establish baseline performance. A range of models was
tested, including traditional machine learning algorithms, deep learning architectures,
and advanced transformer-based models. The results provide valuable insights into how
each model handles the complexity of Telugu offensive language detection, particularly
in code-mixed and low-resource settings.

Table 5.1 Testing Results on Original Dataset

Model Class Precision Recall F1-Score Accuracy

Random
Forest non-hate 0.47 0.86 0.61 0.57

Random
Forest hate 0.81 0.37 0.51 0.57

Logistic
Regression non-hate 0.52 0.78 0.62 0.63

Logistic
Regression hate 0.79 0.53 0.63 0.63

Decision Tree non-hate 0.46 0.76 0.57 0.55

Decision Tree hate 0.72 0.4 0.52 0.55

SVM non-hate 0.53 0.79 0.64 0.64

SVM hate 0.8 0.54 0.64 0.64

Ensemble
Model non-hate 0.46 0.86 0.6 0.55

Dept. of CSE, MVSREC 23


Offensive Language Identification in Telugu

Ensemble
Model hate 0.79 0.34 0.48 0.55

BiLSTM +
Word2Vec non-hate 0.4737 0.0938 0.1565 0.5992

BiLSTM +
Word2Vec hate 0.6099 0.9315 0.7371 0.5992

CNN +
Word2Vec non-hate 0.4375 0.0729 0.125 0.595

CNN +
Word2Vec hate 0.6062 0.9384 0.7366 0.595

CNN +
BiLSTM +
Word2Vec non-hate 0.4375 0.0729 0.125 0.595

CNN +
BiLSTM +
Word2Vec hate 0.6062 0.9384 0.7366 0.595

FastText +
CNN +
BiLSTM non-hate 0.1667 0.0104 0.0196 0.5868

FastText +
CNN +
BiLSTM hate 0.5975 0.9658 0.7382 0.5868

FastText +
BiLSTM non-hate 0.8 0.0417 0.0792 0.6157

FastText +
BiLSTM hate 0.6118 0.9932 0.7572 0.6157

FastText +
CNN non-hate 0.4017 1 0.5731 0.4091

FastText +
CNN hate 1 0.0205 0.0403 0.4091

IndicSBERT non-hate 0.7 0.78 0.74 0.7521

Dept. of CSE, MVSREC 24


Offensive Language Identification in Telugu

IndicSBERT hate 0.93 0.74 0.83 0.7521

mBERT non-hate 0.71 0.72 0.71 0.7727

mBERT hate 1 0.69 0.81 0.7727

IndicBERT non-hate 0.66 0.78 0.71 0.75

IndicBERT hate 0.84 0.73 0.78 0.75

XLM-R non-hate 0.75 0.61 0.67 0.76

XLM-R hate 0.77 0.86 0.82 0.76

From the Table 5.1 of original HOLD dataset results, traditional models such as
Random Forest, Logistic Regression, Decision Tree, and SVM showed varying
accuracy between 55% and 64%. Random Forest had high precision for hate (0.81) but
low recall (0.37), while Logistic Regression offered more balanced performance with
F1-scores of 0.63 for both classes and an accuracy of 63%. SVM achieved the highest
among traditional models with an F1-score of 0.64 for both classes and 64% accuracy.
Decision Tree maintained moderate metrics across both classes with an accuracy of
55%.

Deep learning models like BiLSTM + Word2Vec, CNN + Word2Vec, and CNN
+ BiLSTM + Word2Vec showed high F1-scores for the hate class (around 0.73–0.74)
but extremely low for the non-hate class (as low as 0.12–0.15), indicating class
imbalance. FastText-based combinations performed similarly. In contrast, transformer-
based models outperformed others. IndicSBERT achieved F1-scores of 0.74 (non-hate)
and 0.83 (hate) with 75.21% accuracy. mBERT scored 0.71 and 0.81 for non-hate and
hate respectively, reaching 77.27% accuracy. XLM-R and IndicBERT also showed
high accuracy (76% and 75%) with well-balanced precision and recall across classes.

Table 5.2 Testing Results on Augmented Dataset

Model Class Precision Recall F1-Score Accuracy

non-hate 0.45 0.89 0.59 0.52


Random

Dept. of CSE, MVSREC 25


Offensive Language Identification in Telugu

Forest

Random
Forest hate 0.78 0.27 0.41 0.52

Logistic
Regression non-hate 0.45 0.85 0.59 0.52

Logistic
Regression hate 0.76 0.31 0.44 0.52

SVM non-hate 0.51 0.65 0.57 0.61

SVM hate 0.72 0.59 0.65 0.61

Decision Tree non-hate 0.46 0.76 0.57 0.55

Decision Tree hate 0.72 0.4 0.52 0.55

Ensemble
Model non-hate 0.47 0.83 0.6 0.55

Ensemble
Model hate 0.77 0.37 0.5 0.55

BiLSTM +
Word2Vec non-hate 0.4017 1 0.5731 0.4091

BiLSTM +
Word2Vec hate 1 0.0205 0.0403 0.4091

CNN +
Word2Vec non-hate 0.3983 1 0.5697 0.4008

CNN +
Word2Vec hate 1 0.0068 0.0136 0.4008

CNN +
BiLSTM +
Word2Vec non-hate 0.3975 0.9896 0.5672 0.4008

CNN +
BiLSTM +
Word2Vec hate 0.6667 0.0137 0.0268 0.4008

Dept. of CSE, MVSREC 26


Offensive Language Identification in Telugu

CNN +
BiLSTM +
FastText non-hate 0.4017 1 0.5731 0.4091

CNN +
BiLSTM +
FastText hate 1 0.0205 0.0403 0.4091

FastText +
BiLSTM non-hate 0.4017 1 0.5731 0.4091

FastText +
BiLSTM hate 1 0.0205 0.0403 0.4091

FastText +
CNN non-hate 0.4017 1 0.5731 0.4091

FastText +
CNN hate 1 0.0205 0.0403 0.4091

IndicSBERT non-hate 0.74 0.86 0.8 0.8264

IndicSBERT hate 0.9 0.8 0.85 0.8264

IndicBERT non-hate 0.58 0.81 0.68 0.69

IndicBERT hate 0.83 0.62 0.71 0.69

mBERT non-hate 0.63 0.83 0.72 0.7397

mBERT hate 0.86 0.68 0.76 0.7397

XLM-R non-hate 0.4 1 0.57 0.4

XLM-R hate 0.4 0.69 0.50 0.5

In the augmented dataset results, traditional models like Random Forest and
Logistic Regression showed similar performance trends with reduced overall accuracy
(both at 52%). These models displayed high recall for the non-hate class (0.89 and 0.85
respectively) but poor recall for the hate class (0.27 and 0.31), resulting in lower F1-
scores for hate (0.41 and 0.44). SVM achieved the best performance among traditional
models with an accuracy of 61% and an F1-score of 0.65 for hate, showing more

Dept. of CSE, MVSREC 27


Offensive Language Identification in Telugu

balanced metrics. Decision Tree and the Ensemble Model maintained 55% accuracy,
with modest F1-scores around 0.5 for hate and 0.57–0.6 for non-hate.

Deep learning models such as BiLSTM + Word2Vec, CNN + Word2Vec, and


FastText combinations performed poorly under augmentation. They showed perfect or
near-perfect recall for non-hate (≈1.0) but extremely low recall for hate (as low as
0.0068), resulting in very low F1-scores for the hate class (mostly ≈0.04) and accuracy
stuck around 40.8%. In contrast, transformer-based models again outperformed others.
IndicSBERT showed significant improvement post-augmentation, reaching an
accuracy of 82.64% with F1-scores of 0.8 and 0.85 for non-hate and hate respectively.
mBERT followed with 73.97% accuracy and balanced F1-scores (0.72 and 0.76), while
IndicBERT achieved 69% accuracy. XLM-R’s performance dropped sharply to 50%
accuracy, primarily due to high class imbalance in predictions.

Table 5.3 Comparison of Results for Original and Augmented Datasets

Model Accuracy Avg F1-Score Accuracy Avg F1-Score


(Original) (Original) (Project) (Project)
Random Forest 0.57 0.56 0.52 0.50
Logistic 0.63 0.63 0.52 0.52
Regression
SVM 0.64 0.64 0.61 0.61
Decision Tree 0.55 0.55 0.55 0.55
Ensemble 0.55 0.54 0.55 0.55
Model
BiLSTM + 0.60 0.45 0.41 0.31
Word2Vec
CNN + 0.60 0.43 0.40 0.29
Word2Vec
FastText + 0.62 0.42 0.41 0.31
BiLSTM
IndicSBERT 0.75 0.79 0.83 0.83
mBERT 0.77 0.76 0.74 0.74

Dept. of CSE, MVSREC 28


Offensive Language Identification in Telugu

IndicBERT 0.75 0.75 0.69 0.70


XLM-R 0.76 0.75 0.50 0.54

The comparison shows that while data augmentation led to a slight drop in
performance for traditional and deep learning models, it introduced greater variability
and complexity in the dataset, simulating real-world scenarios more effectively. This
allowed transformer-based models like IndicSBERT to perform significantly better,
achieving the highest accuracy and F1-scores post-augmentation. The results highlight
that augmentation not only stress-tested the models but also revealed which
architectures are more robust and suitable for deployment in noisy, low-resource
language environments like Telugu, marking a meaningful advancement over baseline
approaches.

Dept. of CSE, MVSREC 29


Offensive Language Identification in Telugu

[Link]
In conclusion, this project demonstrates a comprehensive approach to offensive
language detection in Telugu by combining traditional, deep learning, and transformer-
based models with advanced data preprocessing and augmentation techniques. While
some models showed reduced performance post-augmentation, the process enabled a
deeper evaluation of model robustness in handling real-world, noisy, and imbalanced
data. Notably, transformer models like IndicSBERT showed marked improvement,
proving to be more resilient and accurate. This reinforces the importance of data
diversity and contextual understanding in NLP tasks for low-resource languages. The
final system, integrated into a Flask-based application, provides a scalable and practical
solution for real-time offensive content detection, contributing meaningfully to
inclusive and safe digital communication.

Dept. of CSE, MVSREC 30


Offensive Language Identification in Telugu

REFERENCES

Dept. of CSE, MVSREC 31


Offensive Language Identification in Telugu

APPENDIX (SOURCE CODE)


The following appendix contains the source code implemented for offensive language
detection in Telugu. It includes data preprocessing, model training, evaluation, and
web application integration.

#INTERFACE
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-
scale=1.0" />
<title>TØLD</title>

<!-- Favicon -->


<link rel="icon" type="image/png" sizes="16x16" href="/favicon-
[Link]">
<link rel="icon" type="image/png" sizes="32x32" href="/favicon-
[Link]">
<link rel="icon" type="image/x-icon" href="/[Link]">
<link rel="apple-touch-icon" sizes="180x180" href="/apple-touch-
[Link]">
<link rel="manifest" href="/[Link]">

<!-- W3CSS -->


<link rel="stylesheet" href="[Link]
/>
<!-- Font -->
<link rel="stylesheet"
href="[Link]
awesome/5.15.4/css/[Link]">
</head>
<body>
<div id="root"></div>
</body>
</html>

Dept. of CSE, MVSREC 32


Offensive Language Identification in Telugu

#Training the Model


from sklearn.model_selection import train_test_split
Xx_train, Xx_valid, yy_train, yy_valid =
train_test_split(train_df['cleanText'], train_df['enc_label'],
test_size=0.15, random_state=42, stratify =
train_df['enc_label'])
X_train = Xx_train.tolist()
y_train = yy_train.tolist()

X_valid = Xx_valid.tolist()
y_valid = yy_valid.tolist()

X_test = test_with_label['Comments'].tolist()
y_data_with_label = test_with_label['enc_label'].tolist()
from sklearn.linear_model import LogisticRegression

# Logistic Regression
model_lr = LogisticRegression(solver='liblinear', C=1)
model_lr.fit(X_train_tfidf, y_train)
# Predict on the test set
y_pred_lr = model_lr.predict(X_test_tfidf)

# Evaluation
accuracy_lr = accuracy_score(y_data_with_label, y_pred_lr)
classification_report_lr =
classification_report(y_data_with_label, y_pred_lr)
from [Link] import Embedding, Conv1D,
GlobalMaxPooling1D, Dense

Dept. of CSE, MVSREC 33


Offensive Language Identification in Telugu

model = Sequential([
Embedding(vocab_size, 300, weights=[embedding_matrix],
input_length=max_len, trainable=False),
Conv1D(128, 5, activation='relu'),
GlobalMaxPooling1D(),
Dense(1, activation='sigmoid')
])

[Link]()
[Link](optimizer=[Link](learning_rate=0.001),
loss='binary_crossentropy',
metrics=['accuracy'])

history = [Link](
train_pad_sequences,
[Link](train_df['enc_label']),
epochs=100,
batch_size=32,
validation_data=(test_pad_sequences,
[Link](y_data_with_label)),
verbose=1,
callbacks=callback_list
)
]

Dept. of CSE, MVSREC 34


Offensive Language Identification in Telugu

BACKEND
# Flask App Configuration
app = Flask(__name__, static_folder="build", static_url_path="/")
[Link]["SQLALCHEMY_DATABASE_URI"] = "sqlite:///[Link]"
[Link]["SQLALCHEMY_TRACK_MODIFICATIONS"] = False
[Link]["SECRET_KEY"] = "some_secret_key"

# Suppress verbose logs


[Link]('werkzeug').setLevel([Link])
[Link].set_verbosity(0)
tf.get_logger().setLevel('ERROR')

# Initialize extensions
db = SQLAlchemy(app)
CORS(app, supports_credentials=True)
@[Link]("/")
def react_index():
"""Serve React entry point."""
return app.send_static_file("[Link]")
@[Link]("/api/login", methods=["POST"])
def login():
"""Authenticate user and start session."""
try:
email = [Link]("email", "")
pwd = [Link]("pwd", "")
users = getUsers()
user = next(
(u for u in users if [Link](u["email"]) ==
email and [Link](pwd, u["password"])),
None
)

Dept. of CSE, MVSREC 35


Offensive Language Identification in Telugu

if user:
session["user_id"] = str(user["id"])
return jsonify({"success": True})
return jsonify({"error": "Invalid credentials"})
except Exception as e:
print(e)
return jsonify({"error": "Invalid form"})
@[Link]("/api/register", methods=["POST"])
def register():
"""Register a new user."""
try:
email = [Link]("email", "").lower()
pwd_enc = [Link]([Link]("pwd", ""))
username = [Link]("username", "")
if not (email and pwd_enc and username):
return jsonify({"error": "Invalid form"})
if any([Link](u["email"]) == email for u in
getUsers()):
return jsonify({"error": "User already exists"})
if not [Link](r"[\w._]{5,}@\w{3,}\.\w{2,4}", email):
return jsonify({"error": "Invalid email"})
addUser(username, [Link](email), pwd_enc)
return jsonify({"success": True})
except Exception as e:
print(e)
return jsonify({"error": "Invalid form"})

Dept. of CSE, MVSREC 36


Offensive Language Identification in Telugu

@[Link]("/api/logout", methods=["POST"])
@login_required
def logout():
"""Clear user session."""
[Link]()
return jsonify({"success": True})

@[Link]("/api/tweets")
def list_tweets():
"""Return all tweets."""
return jsonify(getTweets())

@[Link]("/api/addtweet", methods=["POST"])
@login_required
def add_tweet():
"""Add a tweet with offensiveness check."""
try:
title = [Link]("title", "").strip()
raw = [Link]("content", "").strip()
clean = [Link](r'<[^>]+>', '', raw).strip()
translit = transliterate_roman_to_telugu(clean)
if not (title and clean):
return jsonify({"error": "Invalid form"}), 422
votes = model_bank.predict(translit)
is_off = votes >= 5
uid = [Link]("user_id")
if addTweet(title, raw, uid, is_off):
return jsonify({"success": True})
return jsonify({"error": "Could not add tweet"}), 400
except Exception as e:
print(e)
return jsonify({"error": "Invalid form"}), 422

Dept. of CSE, MVSREC 37

You might also like