TOXIC COMMENT
DETECTION USING LSTM
A PROJECT REPORT
Submitted by
AKASH KUMAR P 710123243004
GEORGE KIPSON A 710123243017
HARIHARAN R 710123243023
RAGUL S 710123243049
BACHELOR OF TECHNOLOGY
in
DEPARTMENT OF ARTIFICIAL INTELLIGENCE AND DATA SCIENCE
ADITHYA INSTITUTE OF TECHNOLOGY
(Autonomous)
ANNA UNIVERSITY :: CHENNAI - 600 025
NOV / DEC 2025
ADITHYA INSTITUTE OF TECHNOLOGY
(Autonomous)
BONAFIDE CERTIFICATE
Certified that this project report “TOXIC COMMENT DETECTION USING LSTM”
is the Bonafide work of “AKASH KUMAR P (710123243004), GEORGE KIPSON A
(710123243017), HARIHARAN R (710123243023), RAGUL S (710123243049)” who
carried out the project work under my supervision.
PROJECT GUIDE HEAD OF THE DEPARTMENT
Ms. R. Rajalakshmi Mrs. G. Nithya
Assistant Professor, Assistant Professor & Head
Department of AI&DS, Department of AI&DS,
Adithya Institute of Technology, Adithya Institute of Technology,
Coimbatore – 641 107. Coimbatore – 641 107.
INTERNAL EXAMINER EXTERNAL EXAMINER
ACKNOWLEDGEMENT
First of all, we extend our heart-felt Gratitude to the management of
Adithya Institute of Technology, for providing us with all sorts of support in completion
of this Mini Project.
We record our indebtedness to our Director Dr. Joseph V Thanikal, and
our Principal Dr. [Link], for their guidance and sustained
encouragement for the successful completion of this mini project.
We are highly grateful to Dr. [Link], Professor & Dean/
Academic, for his valuable suggestions and guidance throughout the course of this
project, his positive approach had offered incessant help in all possible ways from the
beginning.
We are profoundly grateful to Mrs. [Link], Assistant Professor &
Head, Department of Artificial Intelligence and Data Science for her consistent
encouragement and directions to improve our mini project and completing the mini
project work in time.
We take immense pleasure in expressing our humble note of gratitude to
our project guide, [Link], Assistant Professor, Department of Artificial
Intelligence and Data Science, for her remarkable guidance and useful suggestions,
which helped us in completing the mini project work in time.
We also extend our thanks to other faculty members, Parents and our
Friends for their moral support in helping us to successfully complete the project.
ABSTRACT
Toxic Comment Detection Using LSTM
In the era of digital communication, online platforms have become essential spaces for
interaction, yet they are increasingly plagued by toxic and harmful comments that hinder
healthy discourse. Manual moderation of such content is often inefficient, subjective,
and unable to scale with the growing volume of user-generated data. This project
proposes an automated and intelligent solution for identifying toxic comments through
a Long Short-Term Memory (LSTM)-based Deep Learning architecture. The system is
structured into three primary, interdependent modules. The Data Processing Module
performs text cleaning, tokenization, and embedding to convert unstructured user
comments into numerical sequences suitable for deep learning models. The Training
Module leverages the LSTM’s ability to capture long-range contextual dependencies
within textual data, allowing it to effectively distinguish between toxic and non-toxic
linguistic patterns, even in complex sentence structures. Finally, the Prediction Module
applies the trained model for real-time toxicity classification, outputting a clear toxicity
label, confidence score, and corresponding sentiment analysis. This modular pipeline
highlights the potential of Natural Language Processing (NLP) and Recurrent Neural
Networks (RNNs) in promoting safer online communication environments by offering
a scalable, accurate, and automated comment moderation system.
TABLE OF CONTENTS
Chapter Title Page No.
1 Problem Statement 1
2 Empathy 3
3 Define 8
4 Ideate 12
5 Prototype 17
6 Testing 23
CHAPTER 1
PROBLEM STATEMENT
Social media and online communication platforms play a vital role in the digital
development and social interaction of modern societies, especially for communities that
depend on the internet as their primary means of connection and expression. One of the
major challenges faced by these platforms today is the rapid spread of toxic and offensive
comments, which can severely affect user experience, mental well-being, and community
integrity. A large portion of these harmful messages begin as subtle forms of negativity,
which, if ignored, can quickly escalate into online harassment, cyberbullying, or hate
speech across entire digital communities.
Traditionally, detection of such toxic comments relies on manual moderation by human
reviewers or simple rule-based systems. This approach, although somewhat effective, is
highly time-consuming, subjective, and limited by human consistency. On large-scale
platforms and discussion forums, immediate expert moderation is often unavailable,
resulting in significant delays in identifying and managing harmful content. Consequently,
online communities experience reduced engagement and increased user dissatisfaction due
to delayed or inconsistent moderation.
With the advancement of Artificial Intelligence (AI) and Natural Language Processing
(NLP), there is a growing opportunity to automate the process of toxic comment detection.
Among these technologies, Long Short-Term Memory (LSTM) networks have shown
exceptional performance in understanding language patterns and contextual meaning in
sequential text data. By training LSTM models on large datasets containing
both toxic and non-toxic comments, it becomes possible to automatically classify online
messages with high accuracy. However, developing such a system presents challenges such
as sarcasm, spelling variations, multilingual text, and contextual ambiguity, which can
affect model accuracy. Therefore, there is a need for a robust, scalable, and efficient
framework that can overcome these limitations and provide reliable predictions under real-
world communication scenarios.
The proposed project aims to design and implement an LSTM-based toxic comment
detection system that can automatically identify and categorize different types of online
toxicity from textual data. This system will serve as a decision-support tool for social
media moderators and platform administrators, assisting them in early detection, accurate
classification, and timely action against toxic content. By integrating this model into web
or application-based moderation tools, it can enable widespread adoption and contribute
to creating safer digital environments and more positive user interactions. Online
communication has become the backbone of modern society and is essential for
maintaining healthy global connectivity. However, one of the major threats to digital
platforms is the presence of toxic, hateful, or offensive comments, which directly affect
both the quality of discussions and user trust. Studies indicate that harmful online content
contributes to a significant percentage of digital harassment incidents worldwide.
Traditional moderation techniques rely heavily on manual screening by community
managers or linguistic experts. This process is not only time-consuming and labor-
intensive, but also inconsistent, as accuracy depends on the reviewer’s experience, context
understanding, and available moderation tools.
CHAPTER 2
EMPATHY
Understanding the User
The users of this project are primarily social media moderators, content creators,
researchers, and general internet users who engage in digital communication platforms
such as forums, blogs, and social networks. Online communities today face immense
pressure to maintain respectful and safe interactions. A single toxic comment can escalate
conflicts, harm reputations, and discourage user participation.
Many platforms still rely on manual comment review or keyword-based filtering for
moderation. These methods, while somewhat effective, are time-consuming, inconsistent,
and often fail to capture subtle or context-based toxicity. In smaller or emerging online
communities, access to advanced moderation tools or trained moderators is often limited.
Moreover, varying levels of digital literacy among users make it difficult to interpret
complex moderation policies or AI-based feedback.
Through surveys, interviews, and online observation (real or hypothetical, depending on
the project scope), the following points emerge clearly:
Users often encounter abusive, hateful, or sarcastic comments but struggle to
decide whether to report or ignore them.
Moderators face difficulty identifying disguised toxicity (e.g., sarcasm or coded
hate speech) in large comment sections.
Platform administrators are under constant stress to balance freedom of expression
with maintaining a healthy online environment.
The proposed system should therefore be fast, intelligent, and user-friendly — requiring
minimal technical knowledge.
Understanding the Problem
Online toxicity can spread rapidly and unpredictably across digital platforms, especially
when viral posts trigger emotional discussions. Once a thread or comment section becomes
hostile, regaining civility becomes extremely difficult. Many users fail to recognize subtle
toxicity early on because it often appears as sarcasm, mockery, or emotionally charged
opinions.
For instance:
On social media, a slightly insulting comment can trigger entire chains of hate
replies.
In gaming or discussion forums, toxic chat behavior can quickly evolve into
harassment or group bullying.
On news or video platforms, hate speech and misinformation spread through
emotionally manipulative language.
The lack of effective automated moderation tools leads to:
Delayed intervention, allowing harmful content to influence others.
Inconsistent enforcement of community guidelines.
Emotional distress among users and moderators due to repeated exposure to
abusive content.
Thus, the challenge is not only technical but also social and psychological — to create a
system that understands context, tone, and emotion while promoting a respectful online
culture.
Empathy Mapping
Empathy mapping helps in understanding user emotions, behaviors, and motivations from
different perspectives. It provides a human-centered view of how users and moderators
perceive, feel, and react to online toxicity. Based on observations and user feedback,
people often say they encounter rude or disrespectful comments that make them feel
uncomfortable or unwelcome. Some mention that reporting systems are ineffective or too
slow, while others simply ignore the toxicity due to moderation fatigue.
When we explore what users think, many believe that if they had an intelligent system
capable of detecting and classifying toxicity automatically, it would help maintain
healthier online spaces. Others assume toxic behavior is inevitable in public forums,
reflecting a sense of helplessness and normalization of negativity.
Emotionally, users feel frustrated, anxious, or discouraged when faced with repeated
exposure to toxic content. Moderators often experience stress, burnout, and emotional
fatigue while handling large volumes of harmful posts. Yet, both users and moderators
express hope and trust that AI-driven tools can assist in making digital interactions safer
and more positive.
Need Identification
From the empathy analysis, several essential needs are identified to guide the system’s
design and development:
1. Automation – The system should automatically detect toxic comments from text
data without requiring human intervention.
2. Accuracy – The LSTM model must maintain high precision in differentiating
between toxic and non-toxic comments.
3. Speed – The model should provide real-time predictions for immediate moderation
support.
4. Accessibility – The solution should be deployable on web or mobile platforms with
minimal computational requirements.
5. Usability – The interface must be clear, intuitive, and multilingual, suitable for
users with varying technical skills.
6. Affordability – The tool must be cost-effective and scalable for small to medium-
sized platforms.
7. Awareness – The system should explain the classification (e.g., type of toxicity)
and promote awareness about responsible communication.
Emotional Connection with the Problem
Empathy goes beyond identifying an issue — it involves feeling what the user
experiences. Many users feel emotionally affected when exposed to toxic online
interactions. Moderators and community managers invest time and effort into maintaining
healthy communication spaces, yet often face frustration when toxicity spreads
uncontrollably. Each harmful comment not only affects individual users but also erodes
community trust and inclusivity.
This emotional understanding inspires the creation of a system that not only detects
toxicity but also empowers users with confidence, transparency, and emotional safety.
The goal is to help people trust digital communication again by offering a sense of security
and positivity.
Insights from Empathy
After analyzing user experiences and emotions, several key insights emerge:
Users prefer clear feedback and visual indicators rather than technical reports.
The system should combine simplicity and intelligence — accessible for general
users but powerful for researchers and moderators.
Providing transparency and educational feedback can build user trust in AI
moderation.
Early detection of toxic patterns helps reduce conflicts, stress, and online abuse.
The solution must align with ethical AI principles, promoting fairness, respect,
and inclusivity in online spaces.
Summary
The empathy study highlights that toxic comment detection is not just a technical problem
but a human-centered challenge. By understanding the real struggles, thoughts, and
emotional impact of online toxicity, this project lays the foundation for developing an
effective LSTM-based detection system.
The ultimate goal is to build a bridge between Artificial Intelligence and digital
empathy — where intelligent technology supports respectful communication and
contributes to a safer, more inclusive online world.
CHAPTER 3
DEFINE
After understanding the users’ needs and challenges from the empathy phase, this stage
focuses on defining the main problem and establishing clear design goals for the project
titled “Toxic Comment Detection using LSTM.”
The define phase translates these insights into a structured problem statement that guides
the project’s development. It identifies what needs to be solved, why it is important, and
how the proposed system addresses issues related to online communication and digital
well-being.
In the modern era of social media and digital communication, online toxicity has become
a significant concern. Platforms such as YouTube, Twitter, and Facebook are flooded with
user-generated comments, many of which may contain hate speech, insults, harassment,
or other forms of toxic content.
Traditional moderation systems rely heavily on manual review or simple keyword-based
filters, which are inefficient and often inaccurate. Hence, there is a growing demand for
automated, intelligent systems that can accurately detect and classify toxic comments
using advanced Natural Language Processing (NLP) techniques.
Problem Definition
Online platforms have revolutionized how people communicate, but they also face
challenges from the increasing presence of toxic and abusive comments. Such content
can spread misinformation, promote hate, and discourage healthy online participation.
Manual moderation of comments is time-consuming, inconsistent, and expensive, while
simple rule-based systems often fail to understand context, sarcasm, or linguistic nuances.
For example, a sentence like “You’re unbelievable” can be either positive or sarcastic
depending on the context — something traditional systems struggle to interpret. This leads
to an urgent need for an intelligent, automated, and context-aware system that can
accurately analyze comment sentiment and detect toxic behavior in real time.
Hence, the problem can be defined as follows:
“To develop a deep learning–based system using Long Short-Term Memory (LSTM) networks that can
automatically detect and classify toxic comments from textual input with high accuracy, supporting online platforms
in maintaining positive and respectful digital communication.”
Objectives of the Project
The project sets out several specific objectives to guide its design and implementation:
1. To design and develop an LSTM-based deep learning model capable of
detecting and classifying toxic comments from text data.
2. To preprocess and structure large text datasets effectively using NLP techniques
such as tokenization, stemming, stop-word removal, and word embeddings (e.g.,
Word2Vec or GloVe).
3. To achieve high accuracy, precision, and recall in detecting multiple categories
of toxicity (e.g., threat, insult, hate, or obscene language).
4. To build a user-friendly web application where users can enter comments and
receive instant toxicity analysis and classification.
5. To minimize human moderation effort by automating the identification of
harmful content while maintaining context sensitivity.
Scope of the Project
The scope of this project covers the design, training, testing, and deployment of an
LSTM-based NLP model for toxic comment classification. It primarily foc0uses on four
major aspects — text preprocessing, deep learning model training, performance
evaluation, and deployment.
The project will:
Utilize publicly available datasets such as the Jigsaw Toxic Comment
Classification dataset containing labeled comments across multiple toxicity
categories.
Implement text preprocessing methods including cleaning, tokenization, and
embedding generation for input into the LSTM model.
Develop a web-based application that allows users to input comments and receive
instant predictions about toxicity levels.
Deploy the trained model in a real-time environment suitable for use on websites,
forums, and social media moderation tools.
While the current project focuses on English text comments, future extensions can include
multilingual detection, emotion analysis, and integration with social media APIs for live
moderation.
This project does not focus on audio, video, or image-based toxicity detection, which
can be explored in future research expansions.
Expected Outcomes
Upon successful completion, the project is expected to deliver the following outcomes:
1. High Detection Accuracy: The LSTM model should classify toxic comments with
an accuracy above 90%, ensuring dependable moderation support.
2. Real-Time Toxicity Analysis: The system should provide instant predictions as
users enter comments.
3. Reduced Human Effort: The automation will minimize the need for manual
review by moderators.
4. User-Friendly Interface: The platform should allow easy input and clear
visualization of toxicity scores or labels.
5. Context-Aware Classification: The system should effectively identify subtle toxic
expressions beyond keyword matching.
6. Scalability: The architecture should handle large comment datasets and continuous
data flow from online platforms.
Overall, the expected result is an intelligent, accessible, and socially responsible NLP
system that bridges the gap between artificial intelligence and online safety, empowering
communities and platforms to foster healthier communication.
CHAPTER 4
IDEATE
In the ideation phase, various concepts and potential approaches were explored to
develop an intelligent and reliable system for Toxic Comment Detection using Long
Short-Term Memory (LSTM) networks. The primary objective is to design a deep
learning-based model capable of automatically identifying and classifying toxic
comments from user-generated text on online platforms such as social media, blogs, and
discussion forums.
The process begins with the collection of large-scale datasets containing both toxic and
non-toxic comments from open-source repositories such as the Kaggle Toxic Comment
Classification dataset and other publicly available text corpora. These datasets typically
include categories like hate speech, threats, obscenity, identity attacks, and general toxic
behavior. Before feeding the data to the model, it undergoes a series of preprocessing steps
such as text cleaning, tokenization, stop-word removal, and lemmatization to ensure
consistency and improve training efficiency.
The overall idea is to integrate an accurate, efficient, and scalable model capable of
detecting multiple types of toxicity while maintaining interpretability. Through this
ideation, the project aims to bridge the gap between modern AI-driven Natural Language
Processing (NLP) and practical moderation systems, ensuring that online communication
becomes safer, more respectful, and inclusive.
Model Architecture Ideation
This stage focuses on conceptualizing, comparing, and selecting the most suitable deep
learning architecture for text-based toxicity detection.
1. Brainstorming Model Types
a. LSTM (Long Short-Term Memory):
The foundational idea is to utilize LSTM layers, which excel at understanding long-
range dependencies in textual data. LSTMs can remember the context of earlier
words, making them ideal for detecting toxicity patterns hidden across long
comments or sentences.
b. Bidirectional LSTM (Bi-LSTM):
An extension of LSTM, the Bi-LSTM model reads text sequences in both forward
and backward directions. This approach enables the model to understand contextual
relationships more effectively and improve accuracy.
c. GRU (Gated Recurrent Unit):
GRUs simplify the LSTM structure while retaining performance. They are
computationally efficient and suitable for real-time systems. A GRU-based model
can be considered when hardware or deployment constraints are strict.
d. Hybrid CNN-LSTM Model:
Combining CNN layers with LSTM allows the system to extract local text features
(n-grams or short phrases) through convolutional filters and then capture sequential
dependencies using LSTM layers. This hybrid model can yield higher accuracy in
complex toxicity classification tasks.
2. Selecting the Optimal Model
The final model selection depends on a trade-off between accuracy, computational
efficiency, and deployment feasibility. While transformer models provide state-of-the- art
results, LSTM and Bi-LSTM architectures are chosen for their balance of high
performance, lower resource consumption, and interpretability. This makes them ideal for
scalable deployment across web applications and content moderation systems.
Data Strategy Ideation
Data plays a critical role in determining the accuracy and reliability of the toxic comment
detection system. Hence, an effective data strategy is essential to ensure the dataset is
comprehensive, unbiased, and representative of real-world online communication.
1. Dataset Sourcing
a. Public Datasets:
The Kaggle Toxic Comment Classification Challenge dataset is one of the most
widely used sources. It contains millions of comments labeled under multiple
categories such as “toxic,” “severe toxic,” “obscene,” “threat,” “insult,” and
“identity hate.”
b. Custom Data Collection:
In addition to public datasets, data can be gathered from online platforms such as
Reddit, Twitter, and YouTube comment sections using APIs. Collected data will
be cleaned and manually labeled to include local languages, slang, and contextual
expressions, making the model more robust for multilingual and region-specific
toxicity detection.
2. Data Preprocessing and Augmentation
To enhance the model’s ability to understand text effectively, several preprocessing and
data enhancement steps are ideated:
Text Cleaning: Removing unnecessary symbols, punctuation, numbers, and URLs.
Tokenization: Splitting text into individual words or tokens.
Stop-word Removal: Eliminating common but irrelevant words (e.g., “is,” “the,”
“and”) to focus on meaningful terms.
Lemmatization/Stemming: Reducing words to their base or root forms for uniform
representation.
Word Embeddings: Using embedding techniques like Word2Vec, GloVe, or
FastText to convert words into dense vector representations that preserve semantic
meaning.
Data Augmentation: Introducing paraphrased sentences, synonyms, or back-
translated text to balance class distribution and improve generalization.
Interface and Application Ideation
The platform is designed as a web-based application with a clean and intuitive layout.
Users can input a comment or upload a text file containing multiple comments. Upon
submission, the model processes the input and outputs results such as:
The toxicity classification (e.g., toxic, non-toxic, threat, insult, etc.)
A confidence score showing how certain the model is about its prediction
Additional ideas include:
Real-time Monitoring: Detecting and flagging offensive comments as they are
posted.
Offline Capability: Allowing basic text analysis without continuous internet
connectivity.
Multilingual Support: Handling English and regional languages for broader
inclusivity.
Visualization Dashboard: Providing insights on toxicity trends, frequency, and
category distribution.
CHAPTER 5
PROTOTYPE
Basic Web Application for Comment Input and Toxicity Detection
This is often the first functional prototype, focusing on demonstrating the core LSTM
model’s ability to classify comments as toxic or non-toxic.
• Description:
A simple web interface where a user can type or paste a comment into a text box. The text
is sent to a backend server (e.g., Flask, Django, FastAPI) running the trained LSTM model.
The server processes the text, predicts whether the comment is toxic, and returns the
classification label and confidence score to the user on the web page.
• Key Features:
Text Input: A clean and simple field to type or paste comments.
Prediction Display: Shows whether the comment is “Toxic” or “Non-Toxic,” along
with a confidence percentage.
Category Classification: Optionally classifies comments into multiple categories
such as “Obscene,” “Threat,” “Insult,” or “Identity Hate.”
Minimalistic UI: Prioritizes core functionality for early testing and validation.
• Technologies:
Python (Flask/Django/FastAPI), HTML, CSS, JavaScript, TensorFlow/Keras (for LSTM
model).
Basic Web Application for Comment Input and Toxicity Detection :
Mobile Application (Android/iOS) for Real-Time Comment Filtering
This prototype aims for fast, on-the-go detection of toxic comments in messaging or social
media contexts.
• Description:
A native mobile application that allows users to enter text manually or paste comments
from social media platforms or chat applications. The text is either processed locally (if
using a lightweight model) or sent to a cloud backend running the full LSTM model. The
app returns the toxicity result instantly.
Key Features:
Real-Time Detection: Immediate feedback on comment toxicity while typing.
Offline Capability (Optional): On-device processing with TensorFlow Lite for
quick results without internet.
History Tracking: Stores previously analyzed comments and their toxicity results.
User-Friendly Interface: Simple and responsive design for mobile usage.
Notifications: Warns users when a comment is likely to be toxic before sending or
posting.
• Technologies:
Android Studio (Java/Kotlin) / Xcode (Swift), TensorFlow Lite / Core ML, optional cloud
API backend.
• Use Case:
Helping users avoid posting harmful content and assisting moderators in identifying toxic
language in chats and forums.
Long Short-Term Memory (LSTM) Model
A Long Short-Term Memory (LSTM) network is a special type of Recurrent Neural
Network (RNN) designed to learn long-term dependencies in sequential data, such as text.
Unlike traditional RNNs, LSTM can remember important information over long
sequences and avoid the vanishing gradient problem, making it highly suitable for natural
language processing (NLP) tasks.
In this project, the LSTM model is used to detect whether a comment contains toxic
language. The model processes text as a sequence of words or tokens, learns contextual
meaning, and predicts the level of toxicity. It captures relationships between words—for
example, distinguishing between “not bad” (positive) and “bad” (negative)—which helps
improve classification accuracy.
CHAPTER 6
TESTING
The Testing Module (Performance Assessment)
The Testing Module represents the stage where the trained LSTM model is rigorously
evaluated on a reserved and unseen dataset to determine its true real-world performance in
detecting toxic comments. This ensures that the model not only performs well on training
data but also generalizes effectively to new, unseen text inputs.
A. Core Functions of the Testing Module:
1. Impartial Evaluation:
The testing dataset is strictly separated from both the training and validation datasets
to ensure unbiased performance measurement. The model has never seen this data
before, providing an authentic test of its generalization capability.
2. Prediction Generation:
The trained LSTM model, which performed best on the validation dataset, is used to
process the testing text samples. Each comment in the test set is tokenized, padded
to a fixed sequence length, and passed through the LSTM network. The model then
outputs a predicted class label such as “Toxic,” “Severe Toxic,” “Obscene,”
“Insult,” “Threat,” or “Non-Toxic.”
3. Ground Truth Comparison:
Each prediction generated by the model is compared with its corresponding actual
label from the dataset (Ground Truth). This comparison helps determine how
accurately the model can classify unseen comments.\
4. Error Analysis:
Based on these comparisons, the module records the counts of:
o True Positives (TP): Toxic comments correctly identified as toxic.
o True Negatives (TN): Non-toxic comments correctly identified as non-toxic.
o False Positives (FP): Non-toxic comments incorrectly classified as toxic.
o False Negatives (FN): Toxic comments missed or misclassified as non-toxic.
The Detection Module
The Detection Module is the practical phase where the trained and tested LSTM model is
deployed into a real-world environment for live toxic comment detection. It converts the
deep learning model’s internal computations into actionable, user-facing functionality.
This module operates as the frontline component of the system.
Workflow of the Detection Module:
1. Text Input and Preprocessing:
A user inputs or submits a comment through a web or mobile interface. The text is
immediately sent to the backend for processing.
o Tokenization: The text is split into words or subwords.
o Padding: Shorter sequences are padded to a fixed length to match the model’s
input dimension.
o Encoding: Each token is converted to an integer index based on the vocabulary
used during training.
2. Model Inference (Prediction):
The preprocessed text is then passed through the deployed LSTM model, which
processes the sequence word by word.
The model’s memory gates (input, forget, and output gates) help it interpret context
and detect linguistic patterns such as insults, hate speech, or threats.
3. Classification and Confidence Score:
The model outputs probability scores for each class. The class with the highest
probability is selected as the final label (e.g., “Toxic” with 92.4% confidence).
4. Feedback Interface:
The system displays the classification result and confidence level to the user through
a simple and intuitive interface. In advanced setups, color-coded alerts (e.g., red for
toxic, green for non-toxic) are used for clarity.
This entire pipeline is optimized for speed and efficiency, ensuring that toxic comments
can be flagged instantly — even on low-resource devices. Using frameworks like
TensorFlow Lite or ONNX Runtime, the model can also be deployed on edge devices or
integrated with content moderation APIs.
TRAINING MODULE
The Training Module forms the backbone of the toxic comment detection system. It
involves teaching the LSTM model to understand linguistic structures and semantic
relationships in text that indicate toxicity.
Data Collection and Organization
Dataset Sourcing:
The model is trained using publicly available datasets like the Jigsaw Toxic Comment
Classification Dataset, which contains labeled comments categorized as
toxic, obscene, insult, threat, severe toxic, or identity hate.
Dataset Splitting:
The dataset is divided into three parts:
o Training Set (70–80%) – Used to adjust model weights.
o Validation Set (10–15%) – Used to tune hyperparameters and avoid overfitting.
o Test Set (10–15%) – Used only after training for final performance evaluation.
Example dataset :
Dataset Splitting:
The dataset is divided into three parts:
o Training Set (70–80%) – Used to adjust model weights.
o Validation Set (10–15%) – Used to tune hyperparameters and avoid
overfitting.
o Test Set (10–15%) – Used only after training for final performance evaluation.
o
TRAINING MODEL
The LSTM model is designed with the following key layers:
Embedding Layer: Converts words into dense vector representations capturing
semantic meaning.
LSTM Layer(s): Learns contextual dependencies between words and retains
important information over long sequences.
Dropout Layer: Prevents overfitting by randomly deactivating neurons during
training.
Dense (Fully Connected) Layer: Performs the final classification using sigmoid or
softmax activation functions. The model optimizes its weights using the Adam optimizer
and minimizes binary cross-entropy loss.
Graphs for Training Loss and Accuracy are plotted to monitor progress.
OUTPUT MODULE
The Output Module serves as the user-facing stage, transforming the model’s probabilistic
predictions into clear, understandable results.
After the LSTM processes an input comment and generates probabilities for each class, the
Output Module identifies the most likely category (e.g., “Toxic Comment”) and displays it
in a readable form.
In addition to the classification, the module shows a confidence score to help users gauge
reliability. For example:
Prediction: Toxic Comment (Confidence: 96.3%)
The Output Module may also provide recommendations such as:
“Consider rephrasing this message to make it more polite.”
“This comment contains language that may be considered hateful.”
CONCLUSION
The project “Toxic Comment Detection using LSTM” successfully demonstrates how
deep learning and Natural Language Processing (NLP). By implementing an LSTM-based
architecture, the system effectively learns contextual relationships between words, enabling
it to detect and classify toxic language with high accuracy.
The integration of the Training, Testing, Detection, and Output Modules ensures a
complete, end-to-end solution. The Training Module enables the model to learn from
large labelled datasets, while the Testing Module evaluates its performance and
generalization. The Detection Module allows real-time classification of user comments,
and the Output Module translates the model’s predictions into clear and actionable
insights for users or moderators