ARBA MINCH UNIVERSITY INSTITUTE OF
TECHNOLOGY
FACULTY OF COMPUTING AND SOFTWARE
ENGINEERING
Dep’t: [Link] 2nd year
Program: Regular First semester
Course: Data Warehousing and Data Mining
Project Title: Fake News Classification
By ID
Jilo Dabaso Guyo PRAMIT/104/14
Submitted To: Dr Mohammed Abebe
Submitted date: Feb 24, 2023
Contents
1. Introduction .................................................................................................................................... 1
2. Literature Review ............................................................................................................................ 2
3. Methodology................................................................................................................................... 4
3.1 Method ................................................................................................................................... 4
3.2 Data Gathering and Collection ................................................................................................ 5
3.3. Preprocessing ............................................................................................................................... 5
3.4 Model or Algorithm................................................................................................................. 5
3.4.1 Support Vector Machine ................................................................................................. 5
3.4.2 Logistic Regression .......................................................................................................... 6
3.4.3 Decision Tree Classifier ................................................................................................... 6
3.4.4 KN Neighbours Classifier ................................................................................................. 6
3.4.5 MultinominalNB ..................................................................................................................... 7
4. Dataset Description............................................................................................................................. 7
6. Project Screenshot .......................................................................................................................... 7
6. Experimental Result .......................................................................................................................... 20
7. Conclusion ......................................................................................................................................... 21
Reference .............................................................................................................................................. 22
1. Introduction
Fake news is a type of low quality information, which consists of unethical practices to catch
the attention of readers and listener. It is often spread via traditional media (such as
newspapers) or by posting online (social media). The misinformation passed, even though
marked fake, usually finds its way to thousands and lakhs of people through social media. It
is commonly written to mislead the reader, in order to damage the reputation or image of a
person, entity, company, or a product. Fake news is most commonly detected by headlines in
large font, arrogant language, lavish use of pictures, use of edited and non-existing pictures,
and so on. The frequency of fake news has been observed to increase in the political
hemisphere, especially during the time of elections. Since it is easy to access an online
platform with more than one lakh people, users take advantage to induce public polarization
and create hostile situations by publishing misinformation to gain sympathy and, in the case
of politics, votes [1].
Spreading of misinformation has drastic repercussions, and hence, maximum efforts should
be taken to minimise its influence on civilians, as well as to eradicate the sources of these
fake news. Our primary goal is to break down the different facets of Machine Learning used
for the classification of fake articles as well as arrive at a practical approach to classify real
time data. This might not decrease the scale at which falsities spread, but can act as a medium
to double check the authenticity of any suspicious news posted online.
The project and the paper cover the implementation for predicting the accuracy of a news
article, and also the scope for future work. Thus, the paper consists of the research,
implementation, the result of each approach and potential future advancements in the
following sections: -
Section II: Literature Review - This section presents the technical papers reviewed as a part
of a literature survey, which have played an important role in deepening our understanding. -
Section III: Methodology and Dataset Description - In this section, the architecture of the
system designed is explained, along with the requirement specifications and related diagrams,
and also clear description of dataset described in this section.
Section IV: Experimental results - This section includes the information and pictures of the
implemented tools, the data processing and extracting techniques, and the explanation for the
Data Mining Classification task models used. –
Section V: Conclusion – This section concludes the idea behind the project along with
indication toward future scope of the same.
1
2. Literature Review
Paper 1
WELFake: Word Embedding Over Linguistic Features for Fake News
Detection
Description of the paper: The researchers developed WELFake model that used to classify
news fake or normal. Now days news information are disseminating all over the world
through social media in day to day. The main aim of that researcher is detect fake news. To
successfully did this work the researcher integrate natural language concept with machine
learning algorithm and deep learning algorithm. The first phase pre-processes the data set and
validates the veracity of news content by using linguistic features and then merges the
linguistic feature sets with Word Embedded (WE) and applies machine learning classification
algorithm [2].
Problem Area: Nowadays, social media was very popular for disseminating information
public. However, false information is also shared as true information. One should share
information without caring if it is true. Based on that reasoning, the researcher developed a
model that determines whether the information is fake or true.
What they did? They designed and prepared large dataset to develop fake news detection
model that classify the information as real or fake. They merged four popular news data sets
(i.e., Kaggle, McIntire, Reuters, and Buzz Feed Political) and prepared a more generic data
set of 72134 news articles with 35028 real and 37106 fake news.
They used three stages to develop WELFake model:
They used Linguistic Feature Set for fake news prediction
They used word embedded over LFS for improved detection over WELFake dataset
They Computed Linguistic feature result with state-of-the-art CNN and BERT
methods.
They used CNN and BERT methods for text classification. The state-of-Art CNN used for split
sentence into words (called tokens) and converts them into vectors using context-based WE.
Also the state-of-Art BERT takes input text and pre-processes it using tokenization,
lemmatization, stop word removal, and text lowering operations. Second, it passes the pre-
processed text to the encoding phase where additional token, segment, and positional
embedding processes take place.
Achievement of the paper: Finally, they analysed over 80 linguistic features from state-of-
the-art works and selected 20 significant ones to minimize the computational complexity and
increase the standard classifiers’ accuracy. They applied two WE-based methods (i.e., TF-
IDF, CV) over these linguistic features using six ML models (i.e., KNN, SVM, NB, DT,
Bagging, and AdaBoost) and found out that CV (CountVectorizer) produces better overall
accuracy than TF-IDF with an SVM model. Therefore, they used CV over LFS and classified
the 20 features based on four categories:
2
writing pattern,
readability index,
psycho-linguistics, and
Quantity.
Overall, they Experimental results show that the WELFake model categorizes the news in
real and fake with a 96.73% which improves the overall accuracy by 1.31% compared to
bidirectional encoder representations from transformer (BERT) and 4.25% compared to
convolutional neural network (CNN) models.
Strength of research: They used a large dataset to produce this WELFake model which
helps the model avoid overfitting and also collect additional information to improve the
performance of the model. Researchers briefly explained the background research which
addressed the fake news classification and also clearly defined the gap between their works
with previous research.
Weakness of the Research: The body of the paper is so complex it is difficult to understand.
They used multiple algorithms and multiple techniques in parallel this makes the complexity
of the paper
Suggestion for future research: For future they expected to extend their work with others
for more issues such as knowledge graphs and user loyalty generated model output validation
by WELFake model.
Paper 2
A Robust Chronic Kidney Disease Classifier Using Machine Learning
Description of The paper: the paper explores the clinical support classifier which used to
classify the robust chronic kidney disease that public available to forecast the occurrence of
specific kidney disease chronic. To develop the model they use both biological and historical
data to develop intelligent classifier. They cleaned the data set from outlier and missing, and
by using different techniques determined relevant feature and optimized the model. Finally
they outperformed Support Vector Machine rather than other machine learning models [3].
Problem Area: kidney plays a major role in human health. Chronic kidney disease is the
leading cause of self-inflicted diseases and also many of the world's most worrying diseases.
The researchers developed a model to help with Machine learning technology rather than just
releasing the human impact of the disease to the health professional that is used to detect
chronic kidney disease. Diagnoses based on MRI and CT scan methods are not always
accurate, incur high costs, and are time-consuming. Based on that problem, researchers
propose diagnosis method that needs fewer inputs, which reduces the test cost and prediction time.
Diagnosis method like MRI and CT scan could not available everywhere special at rural areas,
however developed model was more feasible when compared with it.
3
What they did? Researcher acquired the data set from UCI repository. Many activities are
applied on data set to prepare data for the model such as clean up outlier and missing value,
balanced data, select categorical feature, and they used other technique to find the optimal
model design. In the pre-processing stage, these missing data were mainly imputed
using two central tendency values: mean and mode value of respective column. They used
Mean value for numerical features and mode value for categorical features. They used
SMOTE-based oversampling algorithm to balance the dataset in case of preventing
overfitting and under fitting of produced model. They used chi-squared feature selection
approach to build the independent feature correlation that determines the target class.
Researchers used boosting method that refers to hyper parameter tuning which used to define
model design and architecture. Regarding to Hyper Parameter Tuning Grid search was
implemented in order to find the optimal values of the selected HP values.
Achievement of paper: Finally researchers should produce the chronic kidney disease
classifier with 99.33% accuracy produced by Support Vector Machine out of Machine
learning algorithm. They properly prepared data to facilitate dataset for outperform the
model.
Strength of paper: through the model development they focused on the data preparation this
is help them to produce more outperform model with discipline of data preparation. They
applied both parameters which describe independent feature of the dataset and hyper
parameter that used to optimize the model to produce high accuracy mode.
Weakness of paper: the same idea repeated many times in the paper, when notice only brief
concepts are more preferable.
Suggestion for future research: the researcher directs forward that possible future
expansion is with the same or greater accomplishment less feature accuracy values to reduce
the cost of medical testing and more than that. hybrid feature selection techniques can also be
employed to further extract a general and appropriate set of characteristics. Therefore, it can
be estimated using a more robust dataset the severity level of any patient condition.
3. Methodology
A step-by-step and rather meticulous description of the methodology employed in this work
is outlined in the following section.
3.1 Method
When I work this project based up on the following steps or method or process.
Find Dataset Pre- Data Analysis Evaluate
Uses Model
processing models
4
3.2 Data Gathering and Collection
The dataset used for performing the model training in this work was acquired from the
Kaggle. Kaggle is the repository is one of the most reliable and used dataset sources
for researching and implementing machine-learning algorithms. There are 72154 records
and two features in this particular collection, including class attributes such as real and
fake, indicating the type of information.
3.3. Preprocessing
Data preprocessing is a component of data preparation, describes any type of processing
performed on raw data to prepare it for another data processing [Link] has traditionally
been an important preliminary step for the data mining process.
The WELFake dataset comprises missing data that need to be cleaned up during the pre-
processing stage. The dataset containing several missing value such as 565 title records, 57
text records, and 20 label records are missing. We have enough data so simply we dropped
the missing value. In dataset two of columns are text data, so text pre-processing is other step
in the process of building a model.
The Various texts processing step are:
1. Lowercasing: Converting a word to lower case
2. Tokenization: Splitting the sentence into words.
3. Stop words removal: Stop words are very commonly used words (a, an, the, etc.) in
the documents.
4. Stemming: It is a process of transforming a word to its root form.
5. Lemmatization: Unlike stemming, lemmatization reduces the words to a word
existing in the language.
These various text pre-processing steps are widely used for dimensionality reduction.
3.4 Model or Algorithm
A data mining model gets data from a mining structure and then analyses that data by using a
data mining algorithm. The mining structure and mining model are separate objects. The
mining structure stores information that defines the data source. A mining model stores
information derived from statistical processing of the data, such as the patterns found as a
result of analysis.
A mining model is empty until the data provided by the mining structure has been processed
and analysed. After a mining model has been processed, it contains metadata, results, and
bindings back to the mining [Link] this work to develop fake news classification model
five models selected based on the task of project.
3.4.1 Support Vector Machine
Support Vector Machine or SVM is one of the most popular Supervised Learning algorithms,
which is used for Classification as well as Regression problems. However, primarily, it is
5
used for Classification problems in Machine Learning. The goal of the that algorithm is to
create the best line or decision boundary that can segregate n-dimensional space into classes
so that we can easily put the new data point in the correct category in the future. This best
decision boundary is called a hyperplane. The Algorithm chooses the extreme points/vectors
that help in creating the hyperplane.
3.4.2 Logistic Regression
Logistic regression is one of the most popular Machine Learning algorithms, which comes
under the Supervised Learning technique. It is used for predicting the categorical dependent
variable using a given set of independent variables. It predicts the output of a categorical
dependent variable. Therefore the outcome must be a categorical or discrete value. It can be
either Yes or No, 0 or 1, true or False, etc. but instead of giving the exact value as 0 and 1, it
gives the probabilistic values which lie between 0 and 1.
3.4.3 Decision Tree Classifier
Decision Tree is a supervised learning technique that can be used for both classification and
Regression problems, but mostly it is preferred for solving Classification problems. It is a
tree-structured classifier, where internal nodes represent the features of a dataset,
branches represent the decision rules and each leaf node represents the outcome. There
are two nodes, which are the Decision Node and Leaf Node. Decision nodes are used to
make any decision and have multiple branches, whereas Leaf nodes are the output of those
decisions and do not contain any further branches. The decisions or the test are performed on
the basis of features of the given dataset. It is a graphical representation for getting all the
possible solutions to a problem/decision based on given conditions.
It is called a decision tree because, similar to a tree, it starts with the root node, which
expands on further branches and constructs a tree-like structure. In order to build a tree, we
use the CART algorithm, which stands for Classification and Regression Tree algorithm.
A decision tree simply asks a question, and based on the answer (Yes/No), it further split the
tree into subtrees.
3.4.4 KN Neighbours Classifier
K-Nearest Neighbour is one of the simplest Machine Learning algorithms based on
Supervised Learning technique. It assumes the similarity between the new case/data and
available cases and put the new case into the category that is most similar to the available
categories, and used to stores all the available data and classifies a new data point based on
the similarity. This means when new data appears then it can be easily classified into a well
suite category by using K- NN algorithm.
6
K-NN algorithm can be used for Regression as well as for Classification but mostly it is used
for the Classification problems. It is a non-parametric algorithm, which means it does not
make any assumption on underlying data. It is also called a lazy learner algorithm because
it does not learn from the training set immediately instead it stores the dataset and at the time
of classification, it performs an action on the dataset. An algorithm at the training phase just
stores the dataset and when it gets new data, and then it classifies that data into a category
that is much similar to the new data.
3.4.5 MultinominalNB
Multinomial Naive Bayes algorithm is a probabilistic learning method that is mostly used in
Natural Language Processing (NLP). The algorithm is based on the Bayes theorem and
predicts the tag of a text such as a piece of email or newspaper article. It calculates the
probability of each tag for a given sample and then gives the tag with the highest probability
as output.
Naive Bayes classifier is a collection of many algorithms where all the algorithms share one
common principle, and that is each feature being classified is not related to any other feature.
The presence or absence of a feature does not affect the presence or absence of the other
feature.
4. Dataset Description
(WELFake) is a dataset of 72,154 news articles with 37,104 real and 35,028 fake news. For
this, authors merged four popular news datasets (i.e. Kaggle, McIntire, Reuters, BuzzFeed
Political) to prevent over-fitting of classifiers and to provide more text data for better ML
training. Dataset contains four columns: Serial number (starting from 0); Title (about the text
news heading); Text (about the news content); and Label (0 = fake and 1 = real).
6. Project Screenshot
7
8
9
10
11
12
13
14
15
16
17
18
19
6. Experimental Result
Using the dataset discussed above, we trained our model by pre-processing the data and
applying the Natural Language Processing techniques discussed above and finally, used the
20
transformed data as input to the Data Mining models. The accuracy results of classification
are as follows:
Model Accuracy Error
Support Vector Machine 0.94 0.056
Logistic Regression 0.91 0.092
Decision Tree Classifier 0.91 0.075
KNeighbors Classifier 0.84 0.157
MultinominalNB Classifier 0.84 0.164
7. Conclusion
The project undertaken provides a comprehensive and detailed approach to the Data Mining
paradigm towards the fake news classification. This project also provides a broad
understanding of how Data Mining can be used as a tool to combat fake news. For the
uninitiated, this project covers everything, from dataset accumulation, to pre-processing,
tokenization, Stemming, Lemmatization, and vectorising and finally building accurate
classifiers. The Support Vector Machine model gives pretty good results with accuracy
94.0% and 5.6% error rate.
21
Reference
[1] B. A. 2. ,. P. B. 3. R. B. 4. Avinash Bharadwaj 1, “Source Based Fake News Classification using
machine learning,” IJIRSET, 2020.
[2] P. A. I. A. a. R. P. P. K. Verma, “WELFake: Word Embedding Over Linguistic Features for fake news
Detection,” in IEEE Transactions on Computational Social Systems, vol. 8, no. doi:
10.1109/TCSS.2021.3068519., pp. 881-893, 2021.
[3] U. M. B. P. P. M. Debabrata Swain, “A Robust Chronic Kidney Disease Classifier Using Machine
Learning,” electronics, vol. 12, no. 1, 2023.
22