Email Spam Detection Using Naive Bayes in R
A Report submitted under Project Based Learning
In Partial Fulfillment of the Course Requirements for
“R PROGRAMMING” (22CS105001)
Submitted By:
K BHANU PRAKASH REDDY 22102A040532
K R YOGESH 22102A040535
C MAHESH 22102A040537
V YOGESH 22102A040550
G HARI PRASAD 22102A040552
S SAI NARAYANA 22102A040556
Under the Guidance of
Dr. Cuddapah Anitha
Associate Professor Department of CSE, Mohan Babu University
Department of Computer Science and Engineering
School of Computing
MOHAN BABU UNIVERSITY
Sree Sainath Nagar, Tirupati – 517 102 2024-2025
MOHAN BABU UNIVERSITY
Vision
To be a globally respected institution with an innovative and entrepreneurial
culture that offers transformative education to advance sustainability and
societal good.
Mission
❖ Develop industry-focused professionals with a global perspective.
❖ Offer academic programs that provide transformative learning
experience founded on the spirit of curiosity, innovation, and integrity.
❖ Create confluence of research, innovation, and ideation to bring about
sustainable and socially relevant enterprises.
❖ Uphold high standards of professional ethics leading to harmonious
relationship with environment and society.
SCHOOL OF COMPUTING
Vision
To lead the advancement of computer science research and education that
has real-world impact and to push the frontiers of innovation in the field.
Mission
❖ Instil within our students fundamental computing knowledge, a broad
set of skills, and an inquisitive attitude to create innovative solutions
to serve industry and community.
❖ Provide an experience par excellence with our state-of-the-art
research, innovation, and incubation ecosystem to realise our learners’
fullest potential.
❖ Impart continued education and research support to working
professionals in the computing domain to enhance their expertise in
the cutting-edge technologies.
❖ Inculcate among the computing engineers of tomorrow with a spirit to
solve societal challenges.
DEPARTMENT OF COMPUTER SCIENCE AND ENGINEERING
Vision
To become a Centre of Excellence in Computer Science and its emerging areas
by imparting high quality education through teaching, training and research.
Mission
➢ Imparting quality education in Computer Science and Engineering and
emerging areas of IT industry by disseminating knowledge through
contemporary curriculum, competent faculty and effective teaching -
learning methodologies.
➢ Nurture research, innovation and entrepreneurial skills among faculty and
students to contribute to the needs of industry and society.
➢ Inculcate professional attitude, ethical and social responsibilities for
prospective and promising engineering profession.
➢ Encourage students to engage in life-long learning by creating awareness
of the contemporary developments in Computer Science and Engineering
and its emerging areas.
[Link]. Computer Science and Engineering
PROGRAM EDUCATIONAL OBJECTIVES
After few years of graduation, the graduates of [Link]. CSE will be:
PEO1. Pursuing higher studies in core, specialized or allied areas of Computer
Science, or Management.
PEO2. Employed in reputed Computer and I.T organizations or Government to have
a globally competent professional career in Computer Science and
Engineering domain or be successful Entrepreneurs.
PEO3. Able to demonstrate effective communication, engage in teamwork, exhibit
leadership skills and ethical attitude, and achieve professional
advancement through continuing education.
PROGRAM OUTCOMES
On successful completion of the Program, the graduates of [Link]. CSE
Program will be able to:
PO1. Engineering Knowledge: Apply the knowledge of mathematics, science,
engineering fundamentals, and an engineering specialization to the solution
of complex engineering problems.
PO2. Problem Analysis: Identify, formulate, review research literature, and
analyze complex engineering problems reaching substantiated conclusions
using first principles of mathematics, natural sciences, and engineering
sciences.
PO3. Design/Development of Solutions: Design solutions for complex
engineering problems and design system components or processes that meet
the specified needs with appropriate consideration for the public health and
safety, and the cultural, societal, and environmental considerations.
PO4. Conduct Investigations of Complex Problems: Use research-based
knowledge and research methods including design of experiments, analysis
and interpretation of data, and synthesis of the information to provide valid
conclusions.
PO5. Modern Tool Usage: Create, select, and apply appropriate techniques,
resources, and modern engineering and IT tools including prediction and
modeling to complex engineering activities with an understanding of the
limitations.
PO6. The Engineer and Society: Apply reasoning informed by the contextual
knowledge to assess societal, health, safety, legal and cultural issues and the
consequent responsibilities relevant to the professional engineering practice.
PO7. Environment and Sustainability: Understand the impact of the professional
engineering solutions in societal and environmental contexts, and
demonstrate the knowledge of, and need for sustainable development.
PO8. Ethics: Apply ethical principles and commit to professional ethics and
responsibilities and norms of the engineering practice.
PO9. Individual and Team Work: Function effectively as an individual, and as a
member or leader in diverse teams, and in multidisciplinary settings.
PO10. Communication: Communicate effectively on complex engineering activities
with the engineering community and with society at large, such as, being able
to comprehend and write effective reports and design documentation, make
effective presentations, and give and receive clear instructions.
PO11. Project Management and Finance: Demonstrate knowledge and
understanding of the engineering and management principles and apply these
to one’s own work, as a member and leader in a team, to manage projects
and in multidisciplinary environments.
PO12. Life-long Learning: Recognize the need for, and have the preparation and
ability to engage in independent and life-long learning in the broadest context
of technological change.
PROGRAM SPECIFIC OUTCOMES
On successful completion of the Program, the graduates of B. Tech. (CSE) program
will be able to:
PSO1. Apply knowledge of computer science engineering, Use modern tools,
techniques and technologies for efficient design and development of
computer-based systems for complex engineering problems.
PSO2. Design and deploy networked systems using standards and principles,
evaluate security measures for complex networks, apply procedures and
tools to solve networking issues.
PSO3. Develop intelligent systems by applying adaptive algorithms and
methodologies for solving problems from inter-disciplinary domains.
PSO4. Apply suitable models, tools and techniques to perform data analytics for
effective decision making.
PROGRAM ELECTIVE
Course Code Course Title L T P S C
22CS105001 R PROGRAMMING - 1 2 - 2
Pre-Requisite -
Anti-Requisite -
Co-Requisite -
COURSE DESCRIPTION: Introduction to R, R Programming Structures, Doing Math and Simulation in
R, Creating Graphs, Probability Distributions, correlation and Regression and Random Forests.
COURSE OUTCOMES: After successful completion of this course, the students will be able
to:
CO1. Apply R programming constructs to store and manipulate datasets.
CO2. Develop modules using R programming constructs to solve statistical problems.
CO3. Perceive data models to perform descriptive and inferential statistical analysis to
identify trends, patterns in data.
CO4. Create effective visualization using Histograms, Bar plots, Box plots, Scatter
plots for exploratory data analysis.
CO5. Work independently to solve problems with effective communication.
CO-PO-PSO Mapping Table
Program
Course Program Outcomes Specific
Outcome Outcomes
s PO PO PO PO PO PO PO PO PO PO1 PO1 PO1 PSO PSO PSO PSO
1 2 3 4 5 6 7 8 9 0 1 2 1 2 3 4
CO1 3 - - 2 3 - - - - - - - 3 - - -
CO2 3 1 1 1 3 - - - - - - - 3 - - -
CO3 3 3 2 3 3 - - - - - - - 3 - - -
CO4 3 3 2 3 3 - - - - - - - 3 - - -
CO5 - - - - - - - - 3 3 - - - - - -
Course
Correlat
ion 3 2 2 2 3 - - - 3 3 - - 3 - - -
Mappin
g
Correlation Level: 3- High 2-Medium 1- Low
COURSE CONTENT
Module1: INTRODUCTION TO R (08 Periods)
Introduction, How to run R, R Sessions and Functions, Basic Math, Variables, Data Types,
Vectors, Conclusion, Advanced Data Structures, Data Frames, Lists, Matrices, Arrays,
Classes.
Module2: R PROGRAMMING STRUCTURES (10 Periods)
R Programming Structures, Control Statements, Loops, -Looping Over Nonvector Sets,-If-
Else, Arithmetic and Boolean Operators and values, Default Values for Argument, Return
Values, Deciding Whether to explicitly call return-Returning Complex Objects, Functions
are Objective, No Pointers in R, Recursion, A Quicksort Implementation-Extended Extended
Example: A Binary Search Tree.
Module3 DOING MATH AND SIMULATION IN R (10 Periods)
Doing Math and Simulation in R, Math Function, Extended Example Calculating Probability-
Cumulative Sums and Products-Minima and Maxima-Calculus, Functions Fir Statistical
Distribution, Sorting, Linear Algebra Operation on Vectors and Matrices, Extended
Example: Vector cross Product-Extended Example: Finding Stationary Distribution of
Markov Chains, Set Operation, Input /out put, Accessing the Keyboard and Monitor,
Reading and writer Files.
Module4 GRAPHICS (8 Periods)
Graphics, Creating Graphs, The Workhorse of R Base Graphics, the plot() Function –
Customizing Graphs, Saving Graphs to Files.
Module5 PROBABILITY DISTRIBUTIONS AND REGRESSION (9 Periods)
MODELS
Probability Distributions, Normal Distribution-Binomial Distribution-Poisson Distributions
Other Distribution, Basic Statistics, Correlation and Covariance, T -Tests,-ANOVA. Linear
Models, Simple Linear Regression, -Multiple Regression Generalized Linear Models, Logistic
Regression, -Poisson Regression-other Generalized Linear Models-Survival Analysis,
Nonlinear Models, Splines-Decision-Random Forests.
TotalPeriods:45
EXPERIENTIAL LEARNING:
Datatypes, Variables, Operators, Data structures – Vectors, Arrays, Matrices, Lists, Data
frames; Object oriented programming – S3, S4 classes; Selection statements – if statement,
if else statement, switch statement; Iterative statements – For loop, While loop, Repeat
loop, Nested loops; Functions – Creating functions, Default values for arguments, Return
values, Environment and scope issues, Recursion.
1. Create the vectors:
a) (1, 2, 3, . . . , 19, 20)
b) (20, 19, . . . , 2, 1)
c) (1, 2, 3, . . . , 19, 20, 19, 18, . . . , 2, 1)
d) (4, 6, 3) and assign it to the name tmp.
For parts (e), (f) and (g) look at the help for the function rep.
e) (4, 6, 3, 4, 6, 3, . . . , 4, 6, 3) where there are 10 occurrences of 4.
f) (4, 6, 3, 4, 6, 3, . . . , 4, 6, 3, 4) where there are 11 occurrences of 4, 10
occurrences of 6 and 10 occurrences of 3.
g) (4, 4, . . . , 4, 6, 6, . . . , 6, 3, 3, . . . , 3) where there are 10 occurrences
of 4, 20 occurrences of 6 and 30 occurrences of 3.
2. a) Write R code that will generate a vector with the following elements.
"aa" "ba" "ca" "da" "ea" "ab" "bb" "cb" "db" "eb" "ac" "bc" "cc" "dc"
"ec" "ad" "bd" "cd" "dd" "ed" "ae" "be" "ce" "de" "ee"
b) Write a R program to create a Dataframes which contain details of 5
employees and display summary of the data.
3. a) Create a vector of a data set and treat it as an object. Using the vector and
object perform (.) dot product and (x) cross product. Take your own data.
b) “Fizzbuzz” is a simple programming challenge often used at interviews to
test very basic programming skill. Your goal is the following: for the
numbers 1 to 100, print “fizz” if the number is a multiple of 3, “buzz” if the
number is a multiple of 5, “fizzbuzz” if the number is a multiple of both 3
and 5, and simply print the number otherwise.
4. a) Imagine a high school with 1000 lockers all in a row, numbered 1 to 1000
in order. At the start, all of them are closed. 1000 students are sent, one
after the other, to change the state of a set of lockers (from open to closed
or closed to open). The first student changes the state of all lockers. The
second changes the state of every other one (2, 4, 6, 8, . . .
). The third changes the state of every third one (3, 6, 9, 12, . . . ). This
process continues until all 1000 students have gone. Write a R program to
determine which lockers are open at the end of this process?
b) Write a function chomp() that, given a string, removes from the string any
occurrence of the character &, as well as the character to the left of each &
character. So, for example, your function should return:
> chomp ( " a&c " )
"c"
> chomp ( " a&" )
""
> chomp ( " abc " )
" abc "
5. a) Write a function which takes a single argument which is a matrix. The
function should return a matrix which is the same as the function argument
but every odd number is doubled.
b) Write a function that takes an array of numbers x and returns the
smallest number in the array.
Importance and applications of statistical learning, Types of data, Types of
variables, Frequency distributions, Measures of center – Mean, Median, Mode;
Measures of spread – Range, Percentile, Quartiles & Interquartile range,
Standard deviation, Variance; Correlation and Covariance.
6. a) Compute descriptive statistics for the data given below.
X: 14, 20, 22, 19, 15, 18, 30, 27
Y: 16, 25, 27, 20, 16, 18, 27, 23
b) Write a R script which will compute the mean and variance of the vector x
<- 1:100. Compare with R’s internal mean() and var() functions.
7. Write a function to compute running medians. Running medians are a simple
smoothing method usually applied to time-series. For example, for the numbers
7,5, 2, 8, 5, 5, 9, 4, 7, 8, the running medians of length 3 are 5, 5, 5, 5, 5, 5, 7,
7. The first running median is the median of the three numbers 7, 5, and 2; the
second running median is the median of 5, 2, and 8; and so on. Your function
should take two arguments: the data (say, x), and the number of observations for
each median (say, length).
8. Write a R program to perform data import/export (.csv, .xlxs) operations using
data frames in R.
9. Write a R program to create bell curve of a random normal distribution.
10. Write a R program to design correlation matrix by choosing appropriate dataset.
Resources
TEXT BOOKS:
1. The Art of R Programming, Norman Matloff, Cengage Learning
2. R for Everyone, Lander, Pearson
REFERENCE BOOKS:
1. Sandip Rakshit, R for Beginners, McGraw Hill, 2017.
2. Seema Acharya, Data analytics using R, McGraw Hill, 2018.
VIDEO LECTURES:
1. [Link]
2. [Link]
3. [Link]
WEB RESOURCES:
1. [Link]
2. [Link]
3. [Link]
Department of Computer Science and Engineering
CERTIFICATE
This is to certify that the Project Entitled
Email Spam Detection Using Naive Bayes in R
Submitted By
K Bhanu Prakash Reddy 22102A040532
K R Yogesh 22102A040535
C Mahesh 22102A040537
Y Yogesh 22102A040550
G Hari Prasad 22102A040552
S Sai Narayana 22102A040556
is the work submitted under Project-Based Learning in Partial Fulfillment of the
Course Requirements for R PROGRAMMING (22CS105001) during 2025-2026.
Supervisor: Head:
[Link] Anitha Dr. G. Sunitha
Associate Professor Professor & Head
Department of CSE Department of CSE
School of Computing School of Computing
Mohan Babu University Mohan Babu University
Tirupati. Tirupati.
ACKNOWLEDGEMENTS
First and foremost, I extend my sincere thanks to Dr. M. Mohan Babu, Chancellor,
for his unwavering support and vision that fosters academic excellence within the
institution.
My gratitude also goes to Mr. Manchu Vishnu, Pro-Chancellor, for creating an
environment that promotes creativity and for his encouragement and commitment
to student success.
I am deeply appreciative of Prof. Nagaraj Ramrao, Vice Chancellor, whose
leadership has created an environment conducive to learning and innovation.
I would like to thank Dr. K. Saradhi, Registrar, for his support in creating an
environment conducive to academic success.
I am also grateful to Dr. G. Sunitha, Head of the Department of Computer Science
and Engineering, for her valuable insights and support.
Finally, I would like to express my deepest appreciation to my project supervisor, Dr.
Cuddapah Anitha, Associate Professor, Department of Computer Science and
Engineering for continuous guidance, encouragement, and expertise throughout this
project.
Thank you all for your support and encouragement.
Table of contents
Chapter Title Page
No. No.
Abstract 1
1 Introduction 2
1.1 Problem Statement 2
1.2 Importance of the Problem 2
1.3 Objectives 2-3
1.4 Scope of the Project 3
2 System Design 4
2.1 Architecture Diagram Descriptions 4
2.2 Module Descriptions 6
2.3 Database Design 6
3 Implementation 7
3.1 Tools and Technologies Used 7
3.2 Front-End Development 8
3.3 Back-End Development 8
3.4 Integration 8
4 Testing, Results and Discussion 9
4.1 Test Cases 9
4.2 Testing Methods 9
4.3 Output Screenshots 10
4.4 Analysis of Results 10
5 Conclusion 10
5.1 Summary of Findings 10
5.2 Future Enhancements 11
6 Appendix 11
6.1 Code Snippets 11-15
ABSTRACT
With the rapid expansion of digital communication, the prevalence of unsolicited and malicious emails—
commonly known as spam—has increased significantly, leading to wasted time, decreased productivity,
and potential security threats. Manually identifying and filtering these messages is impractical, making
automated detection essential.
This project focuses on developing an automated Email Spam Detection System using the Naive
Bayes classification algorithm in R, a probabilistic machine learning approach well-suited for text
classification tasks. The SMS Spam Collection dataset, containing messages labeled as spam or ham
(legitimate), is used for training and evaluation.
The data undergoes a comprehensive preprocessing pipeline, including conversion to lowercase,
removal of numbers and punctuation, elimination of stopwords, and application of stemming to reduce
words to their root forms. The processed text is then represented as a Document-Term Matrix (DTM),
enabling structured numerical analysis. The dataset is divided into training and testing subsets,
allowing the Naive Bayes model to learn and evaluate patterns in word frequencies associated with
spam and non-spam messages.
Experimental results show that the trained model achieves high accuracy in distinguishing between
spam and ham messages, confirming its robustness for real-world text classification. Furthermore, the
system can effectively classify new, unseen messages, providing a reliable and automated solution
for email filtering.
The implementation demonstrates the potential of machine learning in enhancing email security
and efficiency. Using R ensures flexibility, reproducibility, and scalability, paving the way for future
extensions such as integrating linguistic features, handling larger datasets, or incorporating advanced
hybrid algorithms to further improve spam detection accuracy.
1. INTRODUCTION
1.1 Problem Statement
With the rapid growth of email communication, users are increasingly exposed to unsolicited and
harmful messages, commonly known as spam. Spam emails not only clutter inboxes but can also
carry phishing links, malware, or fraudulent content, causing security and productivity issues for
individuals and organizations. Detecting spam efficiently has become a critical task in email
management systems.
Traditional rule-based spam filters often fail to adapt to the continuously evolving patterns of spam
messages. Machine learning approaches, such as the Naive Bayes classifier, offer a more dynamic
solution by learning from historical email data. Naive Bayes is particularly effective because it can
handle large datasets and calculate the probability of an email being spam based on the frequency
of certain words and phrases.
This project aims to implement an Email Spam Detection system using Naive Bayes in R,
where emails are analyzed, preprocessed, and classified as spam or non-spam. The system will
provide a practical approach to improve email security and reduce the burden of manual filtering.
By leveraging statistical learning, it will help demonstrate the effectiveness of Naive Bayes in text
classification tasks.
1.2 Importance of the Problem
Email is one of the most widely used communication channels in both personal and professional
settings. However, spam emails pose significant challenges by wasting time, reducing productivity,
and sometimes delivering malicious content such as phishing attacks or malware. Efficient spam
detection is essential to protect users and organizations from potential security threats.
Traditional manual or rule-based filtering methods are often ineffective against the constantly
evolving tactics of spammers. Implementing an automated spam detection system using machine
learning techniques, like Naive Bayes, ensures more accurate and adaptive filtering. This reduces
the risk of harmful emails reaching users’ inboxes.
By detecting spam effectively, organizations can maintain cleaner communication channels, improve
workflow efficiency, and safeguard sensitive information. The project also highlights the broader
importance of applying statistical and computational techniques to real-world problems,
demonstrating the value of machine learning in enhancing digital security and user experience.
Finally, the project evaluates the model’s performance to ensure reliable spam detection and safer
email communication.
1.3 Objectives
The primary objectives of this project are:
1. Develop a reliable spam detection system: Build a model using
the Naive Bayes algorithm in R to accurately classify emails as
spam or ham (non-spam).
2. Preprocess and analyze email data: Apply text mining
techniques, such as tokenization, stopword removal, and
stemming, to prepare email data for effective model training.
3. Evaluate model performance: Measure the accuracy, precision,
recall, and F1-score of the Naive Bayes classifier to ensure
robustness and reliability.
4. Assist in automated email filtering: Enable email service
providers or users to automatically filter out unwanted spam
messages, improving productivity and security.
5. Provide a framework for further improvement: Establish a
foundation for incorporating advanced machine learning
techniques for more sophisticated spam detection in the future.
1.4 Scope of the Project
The scope of the Email Spam Detection project using Naive Bayes in R is vast and significant in the
context of modern communication systems, where the proliferation of unsolicited and unwanted
emails has become a major concern for individuals and organizations alike. This project primarily
focuses on building a robust system capable of automatically classifying emails into spam and non-
spam categories, thereby assisting users in managing their inbox efficiently and reducing the time
spent manually filtering messages. The system leverages the Naive Bayes algorithm, which is a
probabilistic classifier based on Bayes’ theorem, and is particularly suitable for text classification
tasks due to its simplicity, effectiveness, and ability to handle high-dimensional data. A key
component of the project is the preprocessing of email data, which involves converting all text to
lowercase to maintain uniformity, removing punctuation, numbers, and special characters that do
not contribute to meaningful classification, and eliminating common stopwords that are frequently
used in emails but do not carry significant information for spam detection. Additionally, the process
of stemming or lemmatization is applied to reduce words to their root forms, ensuring that variations
of the same word are treated consistently during model training. Feature extraction is another
critical aspect, where important words or tokens are identified and represented in a suitable format,
such as a term-document matrix, which allows the Naive Bayes model to learn patterns effectively.
Once the model is trained, it is evaluated using various performance metrics including accuracy,
precision, recall, and F1-score, which provide insights into how well the system can distinguish
between spam and legitimate emails. The scope also extends to practical applications, where the
system can be integrated into email clients or server-side filters to automatically segregate
unwanted messages, thereby improving productivity, reducing the risk of phishing attacks, and
enhancing overall cybersecurity. Moreover, the project serves as a foundation for further research
and development, allowing future improvements by incorporating advanced techniques such as
ensemble learning, support vector machines, or deep learning methods to achieve higher accuracy
and adaptability to evolving spam tactics. The scope also encompasses handling diverse datasets
with emails in multiple formats, including plain text and HTML, and accommodating different
languages or multilingual content, which is increasingly relevant in global communication.
2. System Design
2.1Architecture Diagram Descriptions
The architecture of the Email Spam Detection system using Naive Bayes in R consists of several key
components that work together to collect, process, classify, and evaluate emails. The system can
be visualized in a layered format, with each layer performing a specific role.
1. Data Collection Layer:
This is the first layer of the system where raw email data is collected. Emails may be sourced from
publicly available datasets such as the Enron email dataset or spam collections like the
SpamAssassin dataset. The emails include both spam and ham (non-spam) messages. This layer
ensures that the dataset is representative and sufficiently large to train an effective model.
2. Preprocessing Layer:
Once the raw email data is collected, it passes through the preprocessing layer. This layer is critical
to clean and normalize the text. Steps include converting all text to lowercase, removing
punctuation, numbers, special characters, and stopwords that do not contribute to classification.
Tokenization splits the text into individual words or tokens. Stemming or lemmatization reduces
words to their root form, ensuring consistency across similar words. The output of this layer is clean,
structured data ready for feature extraction.
3. Feature Extraction Layer:
In this layer, important features are extracted from the preprocessed emails. A common approach
is to create a Term-Document Matrix (TDM), which represents the frequency of words in each email.
Each word in the vocabulary becomes a feature, and each email becomes a feature vector. This
structured representation allows the Naive Bayes algorithm to understand patterns and relationships
between words and email classes.
4. Model Training Layer (Naive Bayes Classifier):
Here, the Naive Bayes algorithm is applied to the feature vectors to train the classification model.
The model calculates the probability of an email being spam or ham based on the frequencies of
words in the training data. This probabilistic approach allows the model to make predictions even
when encountering unseen words in new emails. The trained model forms the core of the spam
detection system.
5. Classification Layer:
In this layer, new incoming emails are processed through the same preprocessing and feature
extraction steps and then fed into the trained Naive Bayes model. The classifier computes
probabilities for spam and ham classes and assigns the email to the class with the higher
probability. This layer performs the actual email classification.
6. Evaluation Layer:
After classification, the system’s performance is evaluated. Metrics such as accuracy, precision,
recall, and F1-score are calculated to assess the effectiveness of the model.
2.2 Module Descriptions
The Email Spam Detection system is divided into several key modules, each responsible for a specific
part of the spam detection process.
1. Data Collection Module
This module handles the acquisition of email datasets, which may include spam and ham emails.
Common sources are publicly available datasets like the Enron email dataset or SpamAssassin
dataset. The module ensures that the dataset is balanced and contains enough examples for
accurate model training.
2. Data Preprocessing Module
This module performs text normalization tasks such as converting all text to lowercase, removing
punctuation, numbers, HTML tags, and stopwords. It also includes tokenization, which splits the
text into words or tokens, and stemming or lemmatization to reduce words to their root forms.
3. Feature Extraction Module
This module converts the preprocessed emails into a structured format using techniques like Term-
Document Matrix (TDM) or Bag-of-Words representation. Each email is represented as a vector of
word frequencies, allowing the Naive Bayes model to analyze patterns and probabilities.
4. Model Training Module (Naive Bayes Classifier)
This module applies the Naive Bayes algorithm to the feature vectors. It calculates the probability
of each email belonging to the spam or ham class based on the training data. The trained model
learns to classify emails accurately by recognizing word patterns associated with spam or legitimate
messages.
5. Classification Module
This module receives incoming emails, applies preprocessing and feature extraction, and then uses
the trained Naive Bayes model to predict the class of each email. It outputs whether an email is
spam or legitimate based on computed probabilities.
6. Evaluation Module
This module evaluates the trained model using metrics such as accuracy, precision, recall, and
F1-score. It ensures the reliability of the model and identifies areas for improvement.
7. Output and User Interface Module
This module manages the final output, where spam emails are filtered into a separate folder and
ham emails are kept in the inbox. Optionally, it can include a simple user interface to display results
and provide feedback.
2.3 Database Design
The database design for the Email Spam Detection system is structured to store emails, their
features, and classification results efficiently. The main table, Emails, stores raw email data,
including the sender, subject, body, received date, and a label indicating whether it is spam or ham.
This table serves as the primary source of information for training and testing the Naive Bayes
model, ensuring that the system has access to well-organized and labeled email data.
To support model processing, a PreprocessedEmails table is used to store emails after text
preprocessing. This includes steps such as converting text to lowercase, removing punctuation,
numbers, and stopwords, and applying stemming. Storing preprocessed data separately allows
faster feature extraction and reduces redundant preprocessing during model training or
classification.
The FeatureVectors table holds the numerical representation of emails, often as word frequency
vectors or term-document matrices. These vectors are essential for the Naive Bayes algorithm to
calculate the probability of an email being spam or ham. Once the model classifies the emails, the
results are stored in the ClassificationResults table, which includes the predicted label along with
the probabilities for each class. This enables tracking of predictions and later analysis of misclassified
emails.
Additionally, an optional EvaluationMetrics table can store model performance metrics such as
accuracy, precision, recall, and F1-score. This helps in evaluating the effectiveness of the spam
detection system and provides a foundation for improving the model in the future. Overall, the
database design ensures organized storage, efficient processing, and easy retrieval of email data,
features, and results, forming the backbone of the spam detection system.
3. Implementation
3.1 Tools and Technologies Used
The primary tool used for this project is the R programming language, which is well -suited for
statistical analysis, machine learning, and text mining tasks. R provides a variety of packages and
libraries that make it easier to preprocess data, build machine learning models, and evaluate their
performance. The development is carried out in RStudio, an integrated development environment
(IDE) that offers a user-friendly interface, console, and visualization capabilities.
For text processing and feature extraction, packages such as tm and SnowballC are used. The tm
package helps in cleaning and organizing the email text, removing unwanted characters, stopwords,
and punctuation, while SnowballC allows stemming to reduce words to their root forms. These
preprocessing steps are crucial for preparing the email data for the Naive Bayes classifier.
The Naive Bayes algorithm is implemented using the e1071 package in R, which provides functions
for training the model and making predictions. Optional packages like caret can also be used for
cross-validation and evaluating performance metrics such as accuracy, precision, recall, and F1-
score. Publicly available datasets, such as the Enron Email Dataset or SpamAssassin Dataset, are
used to provide labeled spam and ham emails for training and testing the system.
For better visualization and analysis, R’s plotting functions and packages like ggplot2 are employed.
These tools help in representing email statistics, word frequencies, and model performance in
graphs and charts, making it easier to interpret results. Optionally, a database like MySQL or SQLite
can be used to store emails, preprocessed data, and classification results for organized storage and
retrieval.
3.2 Front-End Development
The front-end development of the Email Spam Detection system focuses on creating a user-friendly
interface that allows users to interact with the system easily. Although the core processing and
classification are handled in R, a simple front-end is designed to accept email input, display results,
and visualize model performance. The interface can be developed using basic web technologies like
HTML, CSS, and JavaScript, or through R-specific visualization tools such as Shiny, which allows
interactive web applications directly from R.
The front-end provides an input section where users can either type an email message or upload a
file containing email content. Once submitted, the email is sent to the back-end for preprocessing
and classification using the Naive Bayes model. After processing, the front-end displays the
classification result as “Spam” or “Ham”, along with optional probability scores indicating the
confidence of the prediction.
Visualization is an important part of the front-end, as it helps users understand the data and model
performance. Graphs showing word frequency, spam versus ham distribution, and model metrics
like accuracy, precision, and recall can be displayed using Shiny’s plotting functions or R
packages such as ggplot2. This makes the system not only functional but also informative for users
who want to analyze the email dataset.
Overall, the front-end development ensures that the Email Spam Detection system is interactive,
easy to use, and provides clear outputs and insights. It acts as a bridge between the user and the
underlying machine learning model, making complex processes simple and accessible for end-users.
3.3 Back-End Development
The back-end development of the Email Spam Detection system is the core of the project, where
all the data processing, model training, and email classification take place. The back-end is primarily
implemented in R, which provides robust libraries and packages for text mining, natural language
processing, and machine learning. This layer handles tasks such as cleaning email text, feature
extraction, training the Naive Bayes model, and making predictions on new emails.
The back-end begins with the data preprocessing module, where raw email content is converted to
lowercase, punctuation and stopwords are removed, and stemming is applied to reduce words to
their root forms. After preprocessing, emails are transformed into feature vectors using techniques
such as a Term-Document Matrix, which allows the Naive Bayes algorithm to analyze patterns in
the data efficiently.
The classification module in the back-end applies the trained Naive Bayes model to the feature
vectors of incoming emails. The model calculates the probabilities of an email being spam or ham
and returns the result to the front-end. Additionally, the back-end includes an evaluation module
that measures model performance using metrics like accuracy, precision, recall, and F1-score. This
ensures that the system is reliable and effective in distinguishing spam from legitimate emails.
Optional back-end components include database integration, where raw emails, preprocessed data,
feature vectors, and classification results can be stored for efficient retrieval and future analysis.
The back-end is designed to be scalable and efficient, capable of handling large volumes of emails,
while also providing flexibility for updates and improvements to the model or preprocessing
techniques.
3.4 Integration
The integration phase of the Email Spam Detection system focuses on combining the front-end and
back-end components to create a fully functional application. In this phase, the user interface, which
allows email input and displays results, is connected with the back-end R modules responsible for
preprocessing, feature extraction, and classification using the Naive Bayes model. This ensures a
seamless flow of data from the user input to the processed output.
Goal: Build a model that classifies emails as Spam or Ham (Not Spam) using Naive Bayes Algorithm
in R.
Algorithm Used: Naive Bayes (a probabilistic classifier based on Bayes’ theorem and independence
assumption).
4. Testing, Results and Discussion
4.1 Test Cases
Testing is a crucial phase in the Email Spam Detection project to ensure that the system accurately
classifies emails and functions as expected. Test cases are designed to evaluate different
components of the system, including data preprocessing, feature extraction, model training,
classification, and the integration between the front-end and back-end. Each test case includes the
input email, expected output (spam or ham), actual output, and the result (pass/fail).
Data Preprocessing Test Cases: These test cases verify whether the email text is cleaned correctly.
For example, emails containing uppercase letters, punctuation, numbers, or special characters
should be converted to lowercase, with all unnecessary elements removed. Tokenization and
stemming are also tested to ensure that words are correctly split and reduced to their root forms.
Classification Test Cases: These test cases focus on the accuracy of the Naive Bayes model. Emails
with known labels (spam or ham) are used as input to check whether the model classifies them
correctly. Edge cases, such as emails with mixed content, unusual symbols, or very short text, are
also tested to ensure the model handles different scenarios effectively.
Integration and Output Test Cases: These test cases evaluate the complete system from front-end
input to back-end processing and output display. For example, when a user submits an email via
the interface, the system should correctly preprocess the text, extract features, classify the email,
and display the result along with confidence scores. Additional test cases can include evaluating the
database storage and retrieval of emails and model results to confirm proper integration of all
modules.
4.2 Testing Methods
Testing methods are applied to ensure that the Email Spam Detection system functions correctly,
efficiently, and accurately classifies emails as spam or ham. The testing process involves unit
testing, integration testing, and system testing to validate each module individually and in
combination, ensuring the complete system works as expected. Each method is designed to identify
errors, improve reliability, and evaluate the overall performance of the system .
Unit Testing: This method focuses on testing individual modules, such as data preprocessing,
feature extraction, and classification. For example, preprocessing is tested to ensure emails are
cleaned properly, stopwords are removed, and stemming is correctly applied. Sim ilarly, the Naive
Bayes classifier is tested using sample emails to check if it predicts the correct class. Unit testing
ensures that each component functions correctly in isolation.
Integration Testing: Integration testing verifies that different modules work together seamlessly.
The front-end interface is tested with the back-end R scripts to ensure emails entered by the user
are correctly processed, classified, and displayed. This method ensures smooth data flow between
modules, proper handling of inputs and outputs, and accurate result presentation.
System Testing: System testing evaluates the complete Email Spam Detection system as a whole.
It involves testing the application under various scenarios, including normal emails, spam emails,
emails with unusual content, and bulk emails. Performance metrics such as accuracy, precision,
recall, and F1-score are calculated to measure the effectiveness of the system. System testing
ensures that the final application meets user requirements and is ready for deployment.
4.3 Output Screenshot
4.4 Analysis of Results
The analysis of results is a critical part of evaluating the effectiveness of the Email Spam Detection
system. After training the Naive Bayes model and testing it on a dataset of emails, the predicted
outputs are compared with the actual labels (spam or ham) to determine the model’s accuracy.
Performance metrics such as accuracy, precision, recall, and F1-score are calculated to assess how
well the system can distinguish between spam and legitimate emails.
The results indicate that the Naive Bayes classifier is highly effective for text-based email
classification. High accuracy suggests that the model correctly identifies most emails, while precision
and recall metrics demonstrate the balance between correctly detecting spam emails and minimizing
false positives. The F1-score provides an overall measure of the model’s reliability in handling both
spam and ham emails.
Through analysis, it is observed that preprocessing plays a significant role in improving results.
Removing stopwords, punctuation, and irrelevant characters, along with stemming, ensures that
the features used for training are meaningful. Well-preprocessed data leads to better feature
vectors, allowing the model to learn patterns more effectively and produce accurate predictions.
Additionally, analyzing misclassified emails provides insights into the limitations of the system.
Emails containing mixed content, very short messages, or unusual formatting may occasionally be
classified incorrectly. These insights can guide future improvements, such as incorporating more
sophisticated text processing, expanding the training dataset, or using ensemble methods to
enhance the model’s performance. Overall, the results confirm that the Email Spam Detection
system is effective, practical, and capable of assisting users in managing unwanted emails
efficiently.
[Link]
5.1 Summary of Findings
The Email Spam Detection project using Naive Bayes in R successfully demonstrated the application
of machine learning techniques for classifying emails as spam or ham. The system was able to
preprocess raw email data effectively, removing irrelevant characters, stopwords, and applying
stemming to generate meaningful features. These features, represented through term-document
matrices, enabled the Naive Bayes model to learn patterns and make accurate predictions.
Testing and evaluation revealed that the system achieved high performance metrics. Accuracy,
precision, recall, and F1-score values indicate that the model is reliable and capable of correctly
identifying the majority of spam and legitimate emails. The results also highlight the effectiveness
of the Naive Bayes algorithm for text classification tasks, especially when combined with proper
preprocessing and feature extraction techniques.
Through analysis of the results, it was observed that preprocessing plays a critical role in model
performance. Well-cleaned and normalized email data ensures that the model can detect spam
patterns more efficiently, while misclassified emails were generally found to be either very short,
containing mixed content, or having unusual formats. These insights provide guidance for future
enhancements, such as expanding the dataset or implementing more advanced text processing
techniques.
Overall, the findings confirm that the Email Spam Detection system is practical, effective, and user-
friendly. It demonstrates that a simple probabilistic approach like Naive Bayes, combined with
proper preprocessing and feature extraction, can significantly improve email management and
reduce spam-related issues, providing a strong foundation for further improvements and real-world
deployment.
5.2Future Enhancements
Although the Email Spam Detection system using Naive Bayes has demonstrated effective
classification of emails, there are several opportunities for future enhancements to improve
accuracy, efficiency, and usability. One key improvement could be the incorporation of advanced
machine learning or deep learning techniques, such as Support Vector Machines (SVM), Random
Forests, or Neural Networks. These methods can capture more complex patterns in email content,
handle larger feature spaces, and improve prediction accuracy, particularly for emails with mixed
or ambiguous content.
Another enhancement is the integration of multilingual support. Currently, the system primarily
focuses on emails in English, but real-world email communication often includes multiple languages or
mixed-language content. By incorporating natural language processing techniques for multiple
languages, the system can become more versatile and applicable in a global context. Additionally, the
system can be improved by implementing real-time spam detection. This would allow incoming emails
to be classified immediately upon receipt, providing instant feedback to users. Coupled with a dynamic
feedback mechanism, misclassified emails can be flagged, stored, and used to retrain the model
periodically, ensuring the system adapts to evolving spam tactics over time.
[Link]
6.1Code Snippets
[Link]("tm") # Text mining
[Link]("SnowballC") # Stemming
[Link]("e1071") # Naive Bayes
[Link]("caTools") # Splitting dataset
library(tm)
library(SnowballC)
library(e1071)
library(caTools)
# Load dataset
# Load only the first 2 columns
data <- [Link]("[Link]", stringsAsFactors = FALSE)[, 1:2]
# Check the data
head(data)
# Convert v1 (spam/ham) to factor
data$v1 <- factor(data$v1)
# Create a corpus from the messages (v2 column)
# Safe lowercase function
toSpace <- content_transformer(function(x) iconv(x, to = "UTF-8", sub = ""))
corpus <- VCorpus(VectorSource(data$v2))
corpus <- tm_map(corpus, toSpace)
corpus <- tm_map(corpus, content_transformer(tolower))
corpus <- tm_map(corpus, removeNumbers)
corpus <- tm_map(corpus, removePunctuation)
corpus <- tm_map(corpus, removeWords, stopwords("english"))
corpus <- tm_map(corpus, stemDocument)
corpus <- tm_map(corpus, stripWhitespace)
# Convert text to DTM
dtm <- DocumentTermMatrix(corpus)
# Optional: remove sparse terms (very rare words)
dtm <- removeSparseTerms(dtm, 0.99)
[Link](123) # for reproducibility
split <- [Link](data$v1, SplitRatio = 0.7) # 70% train, 30% test
train_dtm <- dtm[split, ]
test_dtm <- dtm[!split, ]
# Train Naive Bayes classifier
model <- naiveBayes([Link](train_dtm), train_labels)
# Predict on test set
predictions <- predict(model, [Link](test_dtm))
# Confusion matrix
conf_matrix <- table(Predicted = predictions, Actual = test_labels)
print(conf_matrix)
# Accuracy
accuracy <- sum(diag(conf_matrix)) / sum(conf_matrix)
cat("⬛ Model Accuracy:", round(accuracy * 100, 2), "%\n")
# Test new messages
new_emails <- c("You won a free lottery ticket!",
"Meeting rescheduled to tomorrow morning.")
new_corpus <- VCorpus(VectorSource(new_emails))
new_corpus <- tm_map(new_corpus, content_transformer(tolower))
new_corpus <- tm_map(new_corpus, removeNumbers)
new_corpus <- tm_map(new_corpus, removePunctuation)
new_corpus <- tm_map(new_corpus, removeWords, stopwords("english"))
new_corpus <- tm_map(new_corpus, stemDocument)
new_corpus <- tm_map(new_corpus, stripWhitespace)
new_dtm <- DocumentTermMatrix(new_corpus, control = list(dictionary = Terms(train_dtm)))
pred_new <- predict(model, [Link](new_dtm))
# Display predictions
[Link](Email = new_emails, Prediction = pred_new)