International Journal of Scientific Research in Engineering and Management (IJSREM)
Volume: 09 Issue: 09 | Sept - 2025 SJIF Rating: 8.586 ISSN: 2582-3930
Fake Job Post Detection Using Machine Learning
Mr. Vinay Patel G L2 Rudraswamy M S1
2
Assistant Professor, Department of MCA, BIET, Davanagere
1
Student,4th Semester MCA, Department of MCA, BIET, Davanagere
human resource (HR) departments, where one of the
most critical responsibilities is managing employee
ABSTRACT attrition to minimize turnover. Replacing skilled
employees who leave for other companies incurs costs in
With the exponential growth of online job portals, job
the form of hiring expenses and training for the new
seekers are increasingly vulnerable to fraudulent job
employee. Additionally, both tacit and explicit
postings that exploit their personal data and demand
knowledge is lost when an employee departs, and
money under false pretenses. Traditional approaches for
important social relationships may be disrupted. HR
identifying such scams are predominantly manual, rule-
professionals often struggle to articulate their value
based, and inefficient in addressing the rapidly evolving
creation to their organizations, and one of their
tactics of scammers. This project proposes an intelligent,
responsibilities is to enhance HR effectiveness
automated system for detecting fake job postings using
through improved decision-making. Today, there is a
Machine Learning (ML) and Natural Language
growing trend in HR departments to base decisions on
Processing (NLP). The system collects and preprocesses
data. Data-driven decisions can lead to improved
job-related data, extracts significant textual features, and
organizational performance. A common approach
employs classification algorithms to distinguish between
involves machine learning (ML). Machine Learning
real and fake postings with high accuracy. Through
refers to the method of enabling computers to learn from
extensive data analysis and model training, the system
experience. The idea is for an algorithm to learn from
aims to reduce the impact of employment scams by
datasets and improve as it is exposed to new
proactively flagging suspicious job advertisements. The
information. Potentially, ML could be employed in HR
model is designed for scalability, reliability, and
departments to predict employee attrition. Employee
adaptability, making it a valuable addition to modern
turnover plays a significant role, and there are various
recruitment platforms.
factors influencing turnover that could have detrimental
Keywords – Fake job posting Detection,Employment
effects on the organization. Organizations are actively
scam prevention, Logistic regression, Random forest,
attempting to predict employee turnover and utilize this
support vector machine(SVM), job description
information to reduce turnover rates. With high accuracy
analysis, recruitment metadata, automated detection
in predictions, companies can take necessary actions in a
timely manner for employee retention or succession
1. INTRODUCTION planning. The primary objective is to
Employee turnover has emerged as a significant issue for
anticipateemployee attrition, specifically whether an
all companies today, due to its adverse effects on
employee plans to leave or continue with the
workplace productivity and the timely achievement of
organization.
organizational goals.
Currently, companies invest considerable effort into their
© 2025, IJSREM | [Link] | Page 1
International Journal of Scientific Research in Engineering and Management (IJSREM)
Volume: 09 Issue: 09 | Sept - 2025 SJIF Rating: 8.586 ISSN: 2582-3930
2. RELATED WORK small sample sizes or when prior knowledge is
available.
[1] Recent work on detecting fraudulent online
recruitment posts frames the task as supervised text-
[4] Murtagh (1991) provided one of the early
and-metadata classification problem, where lexical
comprehensive explorations of multilayer
cues (e.g., exaggerated compensation, urgent tone),
perceptrons (MLPs) as versatile tools for both
semantic signals from job descriptions, and
classification and regression tasks. His work
posting/user metadata (account age, contact
highlighted the theoretical foundations of MLPs,
channels, domain reputation) are engineered into
their capacity as universal function approximators,
features and learned by models such as Logistic
and the effectiveness of backpropagation for
Regression, SVMs, Random Forests, and gradient-
training nonlinear models. Since then, numerous
boosted trees.
studies have extended these insights, demonstrating
[2] Naïve Bayes has been one of the most widely how MLPs can model complex, high-dimensional
studied probabilistic classifiers due to its simplicity, relationships across domains such as pattern
efficiency, and surprisingly strong performance in recognition, speech processing, and medical
diverse domains. Rish (2001) conducted an diagnosis.
empirical study highlighting its effectiveness and [5] Cunningham and Delany (2007) offered an
limitations across multiple datasets, showing that extensive review of the k-Nearest Neighbour (k-
although the independence assumption is often NN) algorithm, emphasizing its simplicity,
violated in practice, the model still achieves intuitiveness, and effectiveness across a wide range
competitive accuracy in classification tasks, of classification problems. Their work examined
particularly in high- dimensional settings such as key aspects of k-NN, including distance metrics, the
text categorization and spam filtering. Subsequent impact of the choice of kkk, and strategies for
research has built on these findings by exploring handling high-dimensional data.
extensions such as semi-naïve Bayes, tree-
augmented naïve Bayes, and hybrid approaches that [6] Sharma and Kumar (2016) presented a
relax conditional independence while comprehensive survey on decision tree algorithms,
maintaining computational efficiency. emphasizing their role as one of the most widely
[3] Walters (1988) presented an important used and interpretable methods for classification in
discussion on the application of Bayes’s theorem to data mining. Their study reviewed classical
the analysis of binomial random variables, offering algorithms such as ID3, C4.5, CART, and CHAID,
a probabilistic perspective that complements highlighting differences in splitting criteria, pruning
traditional frequentist methods. His work strategies, and handling of continuous versus
emphasized how Bayesian inference can be categorical attributes.
effectively employed in estimating parameters such [7] Dada et al. (2019) provided an in-depth
as success probabilities, particularly in cases with review of machine learning techniques applied to
© 2025, IJSREM | [Link] | Page 2
International Journal of Scientific Research in Engineering and Management (IJSREM)
Volume: 09 Issue: 09 | Sept - 2025 SJIF Rating: 8.586 ISSN: 2582-3930
email spam filtering, outlining the evolution of rather than proactive, as fraudulent posts remain active
approaches from traditional rule-based systems to until they are reported and assessed. By the time
modern learning- based methods. Their survey measures are implemented, many job seekers may have
already become victims of scams. Another prevalent
discussed widely used algorithms such as Naïve
method is rule-based filtering, where job portals apply
Bayes, Support Vector Machines, Decision Trees,
predefined rules and keyword detection to flag
k-Nearest Neighbors, and ensemble techniques,
questionable job listings. These filters examine job
noting their strengths and limitations in handling
descriptions for phrases such as "easy money," "work
the dynamic and adversarial nature of spam. from home with no skills required," or "registration fee
[8] Breiman (2001) introduced Random Forests required." Although this approach can identify some
as a powerful ensemble learning method that scams, it frequently proves ineffective because
combines multiple decision trees to achieve scammers consistently alter their language to evade
improved classification and regression performance. these filters.
His work demonstrated how the technique leverages Additionally, rule-based filtering results in a significant
bootstrap aggregation (bagging) and random feature number of false positives, where genuine job postings
are erroneously flagged as fraudulent.
selection to reduce overfitting, enhance
The current system also lacks intelligent automation,
generalization, and handle high-dimensional data
rendering it ineffective in recognizing sophisticated
effectively.
scams.
It fails to take into account factors such as recruiter
3. Literature Survey
credibility, salary expectations, and employment
Existing System
benefits, which are essential indicators of job fraud.
The current techniques employed for identifying
Problem Statement
fraudulent job postings are predominantly manual, rule-
The emergence of online job portals has significantly
based, and inefficient, which complicates the fight
enhanced the accessibility of job searching; however, it
against the rising prevalence of online employment
has concurrently resulted in a rise in employment scams.
scams. Numerous job portals depend on human
These scams involve deceptive job postings that lure job
moderators, user reports, and basic keyword filtering to
seekers with the promise of non-existent opportunities in
detect and eliminate fake job postings. Nevertheless,
exchange for monetary payments. Numerous job
these techniques are slow, susceptible to errors, and
seekers, particularly recent graduates and those without
ineffective against cunning scammers who continually
employment, become victims of these fraudulent
adapt their strategies.
schemes, suffering losses of both their finances and
A key method utilized in the current system is manual
personal data. Scammers frequently exploit the names of
verification, wherein job platforms engage teams of
reputable companies to create false job listings, which
human reviewers to scrutinize job postings and pinpoint
undermines the credibility of genuine organizations and
fraudulent ones. This procedure is labor-intensive and
misleads job seekers into financial and identity theft.
lacks scalability, as thousands of job listings are
The absence of an efficient automated system to identify
uploaded each day. Furthermore, some platforms rely on
and eliminate such fraudulent job postings leaves
user- reported complaints, allowing job seekers to flag
thousands of individuals vulnerable to job scams on a
dubious job posts. However, this method is reactive
© 2025, IJSREM | [Link] | Page 3
International Journal of Scientific Research in Engineering and Management (IJSREM)
Volume: 09 Issue: 09 | Sept - 2025 SJIF Rating: 8.586 ISSN: 2582-3930
daily basis. “Predicted Fake Job Details” shows the analyzed
4. RESULT dataset with attributes such as IP Address,
Customer Name, Company Name, Location, and
Qualification, along with their corresponding
classification outcomes. Here, the system highlights
details of job entries that are predicted as suspicious
or fake based on the applied machine learning
model.
5. Proposed System
Fig : 4.1 Buttons for selecting data set In order to overcome the shortcomings of current manual
and rule-based methods, the proposed system presents an
automated model for detecting fake job postings,
utilizing Machine Learning (ML) and Natural Language
Processing (NLP). This system is crafted to intelligently
scrutinize job advertisements, recognize fraudulent
patterns, and categorize them as either authentic or
counterfeit with a high degree of precision. By
Fig : 4.2Result. employing sophisticated classification methods, it
guarantees a more secure online job-seeking
atmosphere, safeguarding candidates against scams.
4.1 Buttons for selecting data set.
The proposed system initiates the process by gathering
The interface in the image belongs to a Fake Job
job posting data from a variety of online recruitment
Detection System, and the button displayed is
platforms, encompassing both legitimate and fraudulent
labeled “Select Data Set File.” The purpose of this
entries. This data is subjected to thorough
button is to allow the user to upload or browse for a preprocessing, during which extraneous text is
dataset file that contains job postings or related eliminated, and critical features such as job title,
information which will be used for analysis. Once company information, recruiter contact details, salary
the user clicks this button, the system opens a file range, and job description are extracted. NLP techniques
selection dialog where the dataset (usually in are utilized to examine textual patterns within job
formats like CSV, Excel, or text) can be chosen.
After loading, the system processes the dataset to
extract features, apply machine learning algorithms,
and classify the records as either fake job posting.
4.2 Result.
This screen from the Fake Job Detection System
represents the stage where the results of the
prediction process are displayed. The section titled
© 2025, IJSREM | [Link] | Page 4
International Journal of Scientific Research in Engineering and Management (IJSREM)
Volume: 09 Issue: 09 | Sept - 2025 SJIF Rating: 8.586 ISSN: 2582-3930
descriptions, pinpointing misleading language frequently • Frequently employed preprocessing methods
employed by scammers. include:
o Tokenization: Dividing text into individual words
or phrases.
Additionally, metadata such as recruiter email addresses o Stemming: Converting words to their base
and website links are validated to uncover form (e.g., "running" → "run").
inconsistencies that may suggest fraudulent behavior. o Removing Stopwords: Excluding common terms
System Requirements Specifications Functional like "the," "is," and "a" that do not contribute significant
Requirements meaning.
The system must effectively identify and categorize 3. Feature Extraction
fraudulent job postings utilizing a machine learning • This process transforms textual information into
model that has been trained on an extensive dataset. The numerical features, enabling machine learning models to
following essential functionalities are necessary: analyze it.
a. Creating a graphical user interface (GUI) for the • Common methodologies include:
collection and processing of real-world job posting Term Frequency-Inverse Document Frequency (TF-
datasets. IDF): Assesses the significance of a word within a
document in relation to the entire dataset.
Developing a preprocessing algorithm to transform
unstructured job descriptions into structured formats 4. Testing Dataset
suitable for the features extracted.
• The model undergoes evaluation using previously
a. Implementing a classification model that reliably
unseen data to assess its effectiveness.
differentiates between genuine and fraudulent job
• The dataset utilized in this phase is distinct from
postings.
the training dataset.
b. Integrating a reporting mechanism
5. Classification
that notifies users regarding • Machine learning classifiers (including Logistic
potential fraudulent job Regression, Random Forest, SVM, or Neural Networks)
advertisements determine whether a job posting is genuine or
Architecture Diagram fraudulent.
5. Architecture Overview • If a posting aligns with the traits of known
fraudulent jobs, it is categorized as Fake.
Training Dataset • Conversely, it is classified as Real if it does not.
• A dataset is compiled that includes both authentic 6. Conclusion
and fraudulent job postings. The proposed Fake Job Detection System
• It encompasses information such as job title, successfully leverages the power of Machine Learning
description, company name, contact information, and and Natural Language Processing to automate the
salary. detection of fraudulent job
• Subsequently, the dataset is divided into training advertisements. By analyzing job descriptions,
and testing subsets. recruiter details, and linguistic patterns, the system
2. Preprocessing identifies deceptive content with notable accuracy and
• The textual information in job postings is sanitized minimal human intervention. The classification
and readied for additional analysis. models trained on real-world datasets demonstrate
© 2025, IJSREM | [Link] | Page 5
International Journal of Scientific Research in Engineering and Management (IJSREM)
Volume: 09 Issue: 09 | Sept - 2025 SJIF Rating: 8.586 ISSN: 2582-3930
that automated approaches can significantly outperform approaches and open research problems,‖ Heliyon, vol.
traditional manual and rule-based methods in terms of 5, no. 6, 2019, doi: 10.1016/[Link].2019.e01802.
speed and effectiveness. With its scalable and user- [8][Link],―ST4_Method_Random_Forest
friendly design, this system not only protects job seekers ,‖ Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001, doi:
from falling victim to employment scams but also 10.1017/CBO9781107415324.004.
enhances the credibility and security of online job
platforms. Future improvements may include
incorporating deep learning models,
real-time data feeds, and integration with online
recruitment services for wider deployment.
7. References
[1] B. Alghamdi and F. Alharby, ―An Intelligent
Model for Online Recruitment Fraud Detection,” J. Inf.
Secur., vol. 10, no. 03,
pp. 155–176, 2019, doi:
10.4236/jis.2019.103009.
[2] I. Rish, ―An Empirical Study of the Naïve
Bayes Classifier An empirical study of the naive Bayes
classifier,‖ no. January 2001, pp. 41–46, 2014. [3] D. E.
Walters, ―Bayes’s Theorem and the Analysis of
Binomial Random Variables,‖ Biometrical J., vol. 30,
no. 7, pp. 817–
825, 1988, doi: 10.1002/bimj.4710300710.
[4] F. Murtagh, ―Multilayer perceptrons for
classification and
regression,‖ Neurocomputing, vol. 2, no.
5–6, pp. 183–197,
1991, doi: 10.1016/0925-2312(91)90023-5.
[5] P. Cunningham and S. J. Delany, ―K - Nearest
Neighbour Classifiers,‖ Mult. Classif. Syst., no. May, pp.
1–17, 2007, doi: 10.1016/S0031-3203(00)00099-6.
[6] H. Sharma and S. Kumar, ―A Survey on
Decision Tree Algorithms of Classification in Data
Mining,‖ Int. J. Sci. Res., vol. 5, no. 4, pp. 2094–2097,
2016,
doi: 10.21275/v5i4.nov162954.
[7] E. G. Dada, J. S. Bassi, H. Chiroma, S. M.
Abdulhamid, A. O. Adetunmbi, and O. E. Ajibuwa,
“Machine learning for email spam filtering: review,
© 2025, IJSREM | [Link] | Page 6
International Journal of Scientific Research in Engineering and Management (IJSREM)
Volume: 09 Issue: 09 | Sept - 2025 SJIF Rating: 8.586 ISSN: 2582-3930
© 2025, IJSREM | [Link] | Page 7
International Journal of Scientific Research in Engineering and Management (IJSREM)
Volume: 09 Issue: 09 | Sept - 2025 SJIF Rating: 8.586 ISSN: 2582-3930
© 2025, IJSREM | [Link] | Page 8