0% found this document useful (0 votes)
6 views22 pages

Analyzing Cybercrime on Reddit

This project analyzes cybercrime discussions on Reddit using topic modeling and NLP techniques, focusing on over 300 posts related to various scams. The research identifies common scam types, strategies employed by scammers, and emotional impacts on victims, revealing insights into the dynamics of cybercrime. The findings aim to inform future awareness tools and automated detection systems to mitigate digital fraud.

Uploaded by

munniprincess27
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views22 pages

Analyzing Cybercrime on Reddit

This project analyzes cybercrime discussions on Reddit using topic modeling and NLP techniques, focusing on over 300 posts related to various scams. The research identifies common scam types, strategies employed by scammers, and emotional impacts on victims, revealing insights into the dynamics of cybercrime. The findings aim to inform future awareness tools and automated detection systems to mitigate digital fraud.

Uploaded by

munniprincess27
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

A Topic Modeling Approach to Analyzing

Cybercrime Posts on Reddit

Panuganti Neha
Undergraduate Student, Department of Mathematics and Scientific Computing
Indian Institute of Technology Kanpur, India

Under the guidance of:

Sourya Joyee De
Assistant Professor, Department of Management Sciences
Indian Institute of Technology Kanpur, India

Sharmishtha Mitra
Professor, Department of Mathematics and Statistics
Indian Institute of Technology Kanpur, India

Indian Institute of Technology Kanpur


ACKNOWLEDGEMENTS

We express our gratitude to the Department of Mathematics and Statistics at the Indian
Institute of Technology Kanpur for providing a conducive research environment.

Signature of the Student Signature of the Supervisors


DECLARATION

I/We hereby declare that the work presented in the project report entitled A Topic Modeling
Approach to Analyzing Cybercrime Posts on Reddit” is written by me in my own words and
contains my own or borrowed ideas. At places where ideas and words are borrowed from other
sources, proper references and acknowledgements, as applicable, have been provided.
To the best of my knowledge, this work does not emanate from or resemble work created by
person(s) other than those mentioned and acknowledged herein.

Name and Signature:

Date:

2
Abstract
Cybercrime is a growing issue in today’s digital world, with increasing cases of scams and
online fraud targeting individuals through various platforms. Reddit, being a space where many
users share personal experiences anonymously, serves as a valuable resource for understanding
how these scams are happening in real life. In this project, I collected over 300 posts from
Reddit using its official API, targeting specific subreddits and keywords related to cybercrime.
The goal was to analyze these posts and identify common scam types, scammer strategies,
victim emotions, and behavioral patterns using topic modeling and NLP techniques.
To begin, I cleaned and preprocessed the text data by combining titles and body content of
posts, removing short texts, and applying custom stopword filtering. I then used BERTopic, a
topic modeling library that uses BERT embeddings along with a custom vectorizer to generate
meaningful clusters of discussion. This helped in automatically identifying topics like romance
scams, phishing, sextortion, card fraud, job scams, and more. I also visualized these topics
using pie charts, bar plots, and word clouds to highlight their distribution.
Beyond topic modeling, I also explored deeper patterns using a combination of manual rule-
based labeling and clustering (via KMeans). I identified victim intent (financial, emotional,
both), impact level, reporting behavior, recovery status, and emotions such as regret, sadness,
or helplessness. I further categorized scam steps like how the victim was contacted, the tactics
used (urgency, fake job offers, phishing), and whether the scam succeeded. I also analyzed
victim demographics (e.g., women, elderly, students) and whether they blamed technology or
themselves.
This end-to-end pipeline helped reveal important insights about how different cybercrimes un-
fold and how victims respond. The patterns discovered could help inform future scam aware-
ness tools or automated warning systems. All steps were performed using Python libraries like
BERTopic, sklearn, matplotlib, pandas, and sentence-transformers. This project was carried
out independently as part of my undergraduate research, under the guidance of Prof. Sourya
Joyee De and Prof. Sharmishtha Mitra at IIT Kanpur.

Keywords: Cybercrime, Reddit,API, Topic Modeling, Scams, BERTopic, Natural Language


Processing, Money Laundering, Online Behavior, Phishing

3
1 Introduction
Imagine the early days of the internet when the digital world was a place of endless promise
and new opportunities. As more people came online, however, so did a new kind of trick-
ery—cybercrimes. Much like the swindlers of old, modern scammers have found innovative
ways to take advantage of unsuspecting users, turning simple emails or messages into sophisti-
cated traps.
The digital revolution has transformed how individuals interact, conduct business, and manage
daily activities. However, this increased reliance on digital platforms has been accompanied by
a significant rise in cybercrime incidents. In India, the escalation is particularly alarming:
• In 2014, cybercrime cases surged by 69% compared to the previous year, with 9,622
cases registered under the Information Technology Act and related sections of the Indian
Penal Code [3].
These statistics underscore the urgent need for a comprehensive understanding of cybercriminal
activities and the development of effective counter measures.
Despite the growing body of research on cybercrime, there remains a gap in understanding the
nuanced experiences of victims as shared in online communities. Traditional studies often rely
on structured datasets or official reports, which may not capture the personal narratives and
evolving tactics discussed in real-time on social media platforms. This research aims to bridge
that gap by analyzing Reddit posts to uncover prevalent cybercrime themes, victim experiences,
and emerging scam strategies.
Today, stories of online scams have become all too common. One Reddit post titled “[US]PayPal
phishing email” tells the story of a user who received what looked like an official message from
PayPal, only to later discover it was a cleverly disguised scam. In another post, a user asked,
“How would this scam work?” sparking a discussion that reveals the many layers behind these
digital tricks. These posts serve as snapshots of how cybercriminals operate—using familiar
names and trusted brands to lure victims into a trap.
Moreover, scams on the internet are not limited to financial fraud. Beyond phishing and identity
theft, there are dating scams and sextortion schemes that prey on personal vulnerabilities. For
example, one post titled ”Asian guy/girl from online dating mentors you...” recounts an incident
where an individual, lured by promises of romance and companionship through an online dating
platform, ultimately found themselves entangled in a scheme designed to extract money and
personal information. Such narratives highlight the diverse nature of cybercrime, showing how
scammers adapt their tactics to exploit both financial and emotional weaknesses.
This paper takes a close look at how cybercrimes are discussed on Reddit by employing a
topic modeling approach. We analyzed a substantial collection of posts gathered from vari-
ous communities—including r/Scams, r/phishing, r/CyberSecurity, r/privacy, r/technology, and
r/hacking—to uncover common themes and patterns. Our primary objectives are to:
• RQ1: Identify prevalent types of cybercrime discussed on Reddit.
• RQ2: Scammer’s Perspective - Understand the strategies and tactics employed by cyber-
criminals as reported by users.
• RQ3: Analyze the emotional and psychological impact on victims.

4
• RQ4: Scammed Perspective - Explore patterns in victim demographics and reporting
behaviors.
By exploring these online conversations, our study aims to paint a clear picture of the modern
cybercrime landscape. The insights obtained have the potential to inform targeted awareness
campaigns, educational initiatives, and advanced automated detection systems aimed at miti-
gating the impact of digital fraud.

2 Previous Work
2.1 Cybercrimes and Online Scams
The transformation of crime in the digital age has produced a multifaceted phenomenon where
traditional fraud converges with modern technology to create complex online scams. Early
studies in cybercrime—such as those by Doe et al. (from the comprehensive overview avail-
able at 4) and Miller and Rogers (5)—focused on the technical underpinnings of phishing
attacks and malware distribution. These seminal works laid the groundwork for understanding
digital offenses, documenting the evolution from crude, easily detectable phishing attempts to
sophisticated, meticulously crafted messages that closely mimic official communications, as
explained in resources such as Kaspersky’s guide on phishing (6).
For instance, Johnson et al. (7) anchored their discourse around the structural analysis of phish-
ing emails, illustrating that even minor alterations in language and visual design can dramat-
ically enhance the credibility of fraudulent messages. Similarly, Lee and Chen (8) collected
participant narratives from online forums and social media platforms to explore the broader
socio-emotional dimensions of cybercrime. Their work reveals that scammers increasingly ex-
ploit personal vulnerabilities through dating scams and sextortion schemes, suggesting that the
lure of romance and trust can be as potent as technical deception in defrauding unsuspecting
victims.
Building on these insights, our work applies a topic modeling approach to Reddit discus-
sions on cybercrime. By analyzing posts from communities such as r/Scams, r/phishing, and
r/CyberSecurity, we aim to uncover latent themes that capture both the technical intricacies and
the human impacts of these scams. For example, one illustrative post titled ”Asian guy/girl
from online dating mentors you...” recounts how a victim, enticed by promises of companion-
ship, was ultimately exploited for financial gain—echoing the trends documented by Lee and
Chen.
By synthesizing insights from both technical analyses and victim narratives, our study con-
tributes to the growing body of literature on cybercrime. This integrated perspective is es-
sential not only for advancing academic understanding but also for developing more effective
countermeasures and raising public awareness against the pervasive threat of online scams, as
highlighted by the work on social engineering in Wiley’s resource (9). Now,” Cybercrimes and
Online Scams” can be effectively divided into the following subsections to address specific
methods employed by cybercriminals:
1. Malicious Websites and Links Exploited by Scammers: This subsection delves into how
cybercriminals create deceptive websites and malicious links to lure individuals into di-
vulging sensitive information or downloading harmful software.

5
2. Fraudulent Interactions via Applications and Emails: This part examines the tactics
scammers use within applications and through email communications to execute fraudu-
lent schemes, including impersonation and exploitation of platform vulnerabilities.

2.1.1 Malicious Websites and Links Exploited by Scammers


Cybercriminals frequently employ malicious websites and deceptive links to deceive individu-
als into divulging sensitive information or downloading harmful software. These scams often
involve mimicking legitimate platforms or creating a sense of urgency to manipulate users into
engaging with fraudulent links.
Examples from Posts I have collected from reddit:
• A user reported receiving an email from a Gmail account purporting to be from PayPal.
The email contained an attached .doc file that appeared suspicious. The attachment
claimed to confirm an order, prompting the recipient to check their PayPal account, which
showed no activity. This is a classic example of phishing, where scammers use fake
documents to trick users into revealing personal or financial information.
• A text message impersonating USPS (but written as ”SSUSPS”) asked the recipient to
confirm their ZIP code via a suspicious link. The message included misspellings and
originated from a non-US phone number, clear indications of a phishing attempt aimed
at stealing personal data.

2.1.2 Fraudulent Interactions via Applications and Emails


Scammers exploit applications and emails to execute fraudulent schemes, often leveraging im-
personation tactics or exploiting platform vulnerabilities to gain victims’ trust.
Examples from Posts I have collected from reddit:
• A Reddit user described how a scammer created a fake Instagram account pretending
to be their friend. The scammer requested the victim’s phone number under the guise
of needing help forwarding a code. This technique is commonly used in SIM-swapping
scams or social engineering attacks to gain access to accounts.
• Another user received an email claiming an account was created with British Airways
using their email address but under a different name (”Ivan”). Although the user accessed
the official website directly (avoiding links in the email), they were unable to verify the
legitimacy of the claim, highlighting how scammers can exploit brand names to confuse
and scare victims into taking unnecessary actions.

2.2 NLP and Cybercrime


The application of Natural Language Processing (NLP) techniques has significantly enhanced
the detection and analysis of cybercrime activities. Early research in this domain focused on
employing topic modeling methods to uncover patterns within textual data related to cyber
threats. For instance, Suryotrisongko and Ginardi (1) proposed utilizing topic modeling for Cy-
ber Threat Intelligence (CTI) applications, analyzing hacker forum datasets to enhance threat
recommendations. Similarly, Hossen et al. (2-) applied both supervised and unsupervised
learning techniques, including Latent Dirichlet Allocation (LDA) and Non-negative Matrix

6
Factorization (NMF), to detect cyber threats from hacker forums. These foundational stud-
ies highlighted the effectiveness of topic modeling in identifying emerging vulnerabilities and
exploit techniques discussed within underground hacking communities.
Building upon these methodologies, Moreno-Vera (2) employed LDA to analyze discussions
in underground hacking forums, aiming to detect and classify vulnerability-related conversa-
tions. This approach facilitated the identification of recurring themes and prevalent discussions
concerning vulnerabilities and exploits. Additionally, Pelofske et al.(3) developed a robust
cybersecurity topic classification tool by training multiple machine learning models on data
from platforms like Reddit and StackExchange, demonstrating the utility of NLP in detecting
cybersecurity-related discussions across diverse internet sources.
Drawing inspiration from these prior works, our study adopts BERTopic, an open-source frame-
work proposed by Grootendorst (4), to analyze Reddit discussions pertaining to cybercrime.
BERTopic has been recognized for its ability to generate coherent topics by leveraging transformer-
based embeddings and a class-based TF-IDF procedure. By integrating BERTopic into our
analysis, we aim to uncover nuanced insights into the discourse surrounding cybercrime on
Reddit, providing a comprehensive understanding of the prevalent themes and discussions
within this domain.

2.3 Reddit and Cybercrime Research


Reddit’s pseudo-anonymous nature and diverse user base make it a rich platform for discus-
sions on various topics, including cybercrime. Researchers have tapped into this resource to
analyze user-generated content related to cybersecurity. For example, Silva et al.(5) conducted
sentiment analysis and topic modeling on Reddit comments using VADER and BERTopic, re-
spectively, to identify trends and patterns in user interactions. Their work demonstrated the
effectiveness of these models in classifying emotions and themes expressed in Reddit discus-
sions.
Furthermore, studies have explored the application of advanced topic modeling techniques to
Reddit data. Kaur and Wallace (6) compared unsupervised topic modeling methods, including
BERTopic, to analyze online communities, highlighting the advantages of BERTopic over tra-
ditional approaches like LDA and NMF in capturing nuanced topics within Reddit discussions.
Building upon these methodologies, our study employs BERTopic to analyze Reddit discus-
sions pertaining to cybercrime. This approach enables us to uncover prevalent themes and
patterns within these conversations. We commence with a quantitative approach to identify rel-
evant clusters, followed by an in-depth qualitative validation by researchers. This two-pronged
approach helps map out the primary themes and sub-themes, offering an understanding of the
multifaceted views on cybercrime as discussed on Reddit.

3 Ethics with regards to Reddit Data


In this study, we collected Reddit posts relevant to cybercrime and online scams using publicly
accessible data sources. The keywords used for extraction were carefully selected to reflect
common cybercrime themes, including phishing, fraud, sextortion, dating scams, identity theft,
and money laundering. Posts were gathered from publicly visible subreddits such as r/Scams,
r/CyberSecurity, r/Phishing, r/FraudPrevention, and others related to cyber-
security and scam reporting.

7
To ensure ethical integrity, we followed these key principles:
• Public data only: No data was collected from private or restricted subreddits. All content
was accessed using Reddit’s official API from publicly available threads.
• User anonymity: Personally identifiable information such as usernames and IDs was
not stored, or used in analysis or reporting.
• Dynamic subreddit filtering: Any subreddit that turned private during or after our col-
lection period was excluded from the final dataset.
This project was carried out for academic research purposes and follows standard ethical prac-
tices for analyzing social media data, particularly when handling user-generated discussions on
sensitive topics.

4 Methodology
Topic modeling and computational techniques are widely used for analyzing unstructured datasets,
especially when manual coding is infeasible. This study assumes that topics extracted from
Reddit posts reflect genuine scam strategies and victim experiences shared by users. To ana-
lyze these discussions, we implemented a multi-step methodology that combines data retrieval,
preprocessing, and computational analysis.
As shown in Figure 1, our methodology comprises three stages:
• Data Retrieval & Preprocessing
• Computational Analysis (Topic Modeling, Clustering, and Labeling)
• Labeling and Visualization

Figure 1: Three-step methodology for analyzing Reddit scam-related discussions.

8
4.1 Data Retrieval and Preprocessing
4.1.1 Data Retrieval
Reddit provides an API for accessing posts, which is essential for collecting large-scale textual
data. To access this API, we created an application on the Reddit Developer Portal, obtained
authentication credentials (client ID, client secret, username, password, user agent), and used
Python for data extraction. The complete data collection code is available online.(code)

1) Reddit API Authentication To authenticate, we used the praw library in Python. The
setup is shown below:
import praw

reddit = [Link](
client_id=REDDIT_CLIENT_ID,
client_secret=REDDIT_CLIENT_SECRET,
username=REDDIT_USERNAME,
password=REDDIT_PASSWORD,
user_agent=REDDIT_USER_AGENT
)

2) Keywords and Subreddit Selection The keywords used for extracting scam-related posts
were selected based on their relevance to cybersecurity threats and online fraud. Common
terms included: cybercrime, fraud, phishing, identity theft, romance scam, crypto scam, fake
refund, SIM swap, sextortion, money laundering, and more.
Relevant subreddits included: r/BitcoinScams, r/CryptoScams, r/CyberCrime,
r/OnlineDating, r/ScamReports, r/Scams, r/TechSupportScams, among
others.
All posts were merged into a single dataset for analysis.

3) Data Description The dataset consists of the following columns:


• id – Unique identifier for each post.
• title – The post title.
• text – Main content of the post.
• subreddit – Source subreddit.
• author – Poster’s username.
• score – Number of upvotes.
• num comments – Number of comments.
• url – Direct URL to the post.
• created utc – Timestamp of post creation.

9
4.1.2 Data Cleaning
To ensure robust analysis, the following cleaning steps were applied:
• Handling Missing Data: Empty fields were replaced with empty strings.
• Content Merging: Title and text fields were combined into a single content column.
• Noise Removal: Posts with fewer than 20 characters, or marked as [deleted] or
[removed], were excluded.
• Stopword Removal: Common stopwords were removed using NLTK’s English list and
a custom CountVectorizer.

4.2 Computational Analysis


4.2.1 Topic Modelling by BERTopic- Why BERTopic?
We have decided to choose BERTopic because it effectively combines state-of-the-art transformer-
based embeddings with advanced clustering techniques, yielding coherent topics that capture
the semantic nuances of textual data. This method leverages BERT embeddings, UMAP for
dimensionality reduction, and HDBSCAN for clustering—making it robust and scalable for
our analysis of cybercrime-related posts.

Overview of BERTopic Pipeline


BERTopic transforms a collection of documents into meaningful topics through a multi-step
process:
1. Embedding Generation: Convert documents into high-dimensional vectors using BERT.
2. Dimensionality Reduction: Use UMAP to project these embeddings onto a lower-
dimensional space.
3. Clustering: Apply HDBSCAN to group similar documents based on density.
4. Topic Representation: Extract representative terms for each cluster via a class-based
TF-IDF (c-TF-IDF) approach.

Detailed Mathematical Background


Transformer-based Embeddings
At the core of BERTopic lies a transformer model, such as BERT, which computes contex-
tual embeddings for each document. The transformer leverages the self-attention mechanism,
whose core computation is:

QK T
!
Attention(Q, K, V ) = sof tmax √ V,
dk

where Q, K, and V represent the query, key, and value matrices, respectively, and dk is the
dimension of the key vectors. This mechanism allows the model to weigh the importance of

10
different words in a sentence and to generate a semantic representation Ed for each document
d:
Ed = fBERT (d).
This high-dimensional embedding captures intricate contextual relationships in text, serving as
the foundation for subsequent processing.

Dimensionality Reduction with UMAP


The embeddings {Ed } are high-dimensional and require dimensionality reduction to efficiently
cluster similar documents. UMAP (Uniform Manifold Approximation and Projection) is em-
ployed for this purpose. UMAP builds a weighted k-nearest neighbor graph based on the local
distances between points. It then optimizes a low-dimensional representation by minimizing
the following cross-entropy loss between the high-dimensional and low-dimensional fuzzy sim-
plicial sets: " #
X wij ′
L= wij log ′ + wij − wij ,
(i,j)
wij

where wij are the weights (probabilities) in the high-dimensional space, and wij are the corre-
sponding weights in the low-dimensional embedding. This process yields reduced embeddings:
EUd M AP = U M AP (Ed ),
which preserve the essential topological structure of the data.

Clustering with HDBSCAN


With the low-dimensional embeddings, clustering is performed using HDBSCAN (Hierarchical
Density-Based Spatial Clustering of Applications with Noise). HDBSCAN identifies clusters
by:
• Computing the mutual reachability distance between points, which adjusts the raw
distance based on local density.
• Constructing a minimum spanning tree (MST) from these distances.
• Extracting clusters as dense regions separated by significant gaps in the MST.
Mathematically, for points xi and xj , the mutual reachability distance is defined as:
dmreach (xi , xj ) = max{corek (xi ), corek (xj ), d(xi , xj )},
where corek (x) is the distance from x to its kth nearest neighbor. Clusters are then identified
as sets:
Cluster = {d ∈ D : density(d) ≥ ϵ},
with ϵ being a threshold that separates dense clusters from noise.

Topic Representation using c-TF-IDF


After clustering, each cluster is treated as a “class” and the c-TF-IDF method is used to extract
key terms that characterize the cluster. For a given term t in cluster C, the c-TF-IDF score is
computed as: !
N
c − T F − IDF (t, C) = T F (t, C) × log ,
DF (t)

11
where:
• T F (t, C) is the term frequency of t within cluster C.
• N is the total number of documents.
• DF (t) is the document frequency of t across the entire corpus.
This metric accentuates terms that are frequent in a cluster but rare across the whole corpus,
thereby yielding representative topic descriptors.

Advantages of the Mathematical Approach in BERTopic


The integration of these techniques provides multiple advantages:
1. Semantic Richness: Transformer-based embeddings capture deep contextual relation-
ships beyond simple word counts.
2. Effective Noise Reduction: UMAP and HDBSCAN work together to reduce dimension-
ality while effectively isolating noise, ensuring that only meaningful clusters are formed.
3. Scalable Topic Extraction: The c-TF-IDF approach efficiently extracts topics even from
large datasets by emphasizing cluster-specific terminology.
4. Robustness: The combination of density-based clustering and modern embeddings leads
to stable and coherent topics, even in the presence of outliers.

4.3 Results and Visualization


After applying the BERTopic pipeline, we extracted and analyzed the topics generated from
the dataset. This section outlines the topic extraction process, cluster naming methodology,
and corresponding visualizations.

Step 1: Topic Extraction


Topic Weights and Top Words : We used the get topic info() function from BERTopic
to retrieve the full list of topics. For each topic (excluding topic -1, which includes outliers),
we extracted the top 10 words along with their c-TF-IDF weights. These top words helped us
assign semantic meanings to each cluster.

Cluster Naming (Manually done): Each cluster was assigned a descriptive name based on
its most representative words. The table below shows the interpreted topics, keywords, and
corresponding meanings.

12
Table 1: Topic Interpretations Generated by BERTopic

Topic Top Words Interpretation


0 security, https, cybersecurity, Cybersecurity Events & Data
incident, data, events Breaches
1 money, mom, family, ro- Romance & Family Scams
mance, scammed, pay
2 account, scammer, link, dis- Fake Links & Tech Platform Scams
cord, fake, usps, phone
3 sextortion, bitcoin, email, Email-Based Sextortion & NSFW
videos, nsfw Scams
4 credit, card, debit, charges, Credit/Debit Card Fraud
fraudulent, bank
5 phishing, chrome, extension, Phishing via Browser Extensions
page, client, branded
6 laundering, chinese, myan- Money Laundering & International
mar, aml, money Scams
7 job, phishing, email, scam, Job Offer & Email Phishing
google, pcloud
8 number, phone, walmart, call, Phone Call Scams (Retail & Local)
asked, local

Interpretation of Visualizations
The visualizations in the Figures in the next page provide complementary perspectives topics
formed by bertopic model on our dataset:
• Topic Distribution : Cybersecurity events dominate the discussion (17.8%), followed
by Romance & Family scams (14.6%) and Email-Based Sextortion (13.4%). This dis-
tribution reflects the prevalence of both technical and emotional fraud vectors in online
communities.
• Topic Word Clouds: Each topic exhibits distinct linguistic patterns. Cybersecurity dis-
cussions center around ”security,” ”https,” and ”incident,” while Romance scams fea-
ture terms like ”romance,” ”mom,” and ”family.” Sextortion scams prominently include
”email,” ”content,” and ”bitcoin,” revealing common communication mediums and pay-
ment methods used in these scams.
• Topic Post Count : For each cluster, representative terms are extracted using the for-
mula of c-TF-IDF scores. These justify our manual naming of the clusters based on the
dominance of the words in each cluster/topic formed.
This multi-dimensional analysis highlights the variety and frequency of scam types discussed
across Reddit. The high volume of cybersecurity topics suggests a strong awareness of tech-
nical threats, whereas the substantial presence of romance scams indicates the persistent rele-
vance of emotional manipulation in digital fraud schemes.

13
(a) Topic distribution by percentage

(b) Word clouds for identified topics

14

(c) Post count per topic


Step 2: Similarity Analysis and Topic Relationships
After identifying the nine distinct scam topics through BERTopic, we analyzed how these topics
relate to each other using 2 visualizations: hierarchical clustering and similarity matrix.

Hierarchical Clustering of Topics


The hierarchical clustering dendrogram reveals the semantic relationships between different
scam types based on their shared linguistic patterns. This visualization groups similar topics
together based on their content similarity, creating a tree-like structure that shows:
• A major division between two primary clusters:
– Information-based scam cluster (red branch): Topics 0 (Cybersecurity Events), 5
(Phishing Extensions), and 6 (Money Laundering) form a distinct group focused on
technical and cybersecurity aspects.
– Transaction-based scam cluster (blue branch): Topics including Romance Scams,
Credit Card Fraud, Sextortion, and Platform Scams form another major group cen-
tered around financial transactions and personal exploitation.
• Notable sub-clusters include:
– Topics 1 (Romance Scams) and 4 (Credit Card Fraud) show moderate similarity
( 0.55), indicating overlapping victim narratives.

Figure 3: Hierarchical clustering dendrogram showing semantic proximity between scam top-
ics.

Similarity Matrix Analysis


The similarity matrix heatmap provides a detailed view of cross-topic relationships:
• Topic self-similarities (diagonal): Each topic shows perfect similarity (1.0) with itself,
visualized as dark blue squares along the diagonal.
• High cross-topic similarities:
– Topics 2 (Tech Platform Scams) and 3 (Sextortion) have high similarity ( 0.8), sug-
gesting overlapping tactics.
– Topics 7 (Job Phishing) and 8 (Phone Scams) also show similarity ( 0.7), implying
similar language use.
• Low cross-topic similarities:

15
– Topic 0 (Cybersecurity Events) and Topic 6 (Money Laundering) are linguistically
distinct from most other topics.

Figure 4: Heatmap representing similarity scores between all topic pairs. Darker cells indicate
higher similarity.

These visualizations reveal how different scam types cluster naturally based on their linguistic
patterns. The hierarchical structure highlights that while each scam type has unique character-
istics, there are clear relationship patterns that could inform prevention strategies.
The similarity matrix further quantifies these relationships, showing that while some topics
like Cybersecurity Events maintain distinct boundaries, others like Tech Platform Scams and
Sextortion have considerable overlap, suggesting that victims of one type might be vulnerable
to the other.

Step 3: Analysis of Scammer Intent, Victim Sentiment, Im-


pact, Reporting, and Recovery
The visualizations provide insights into the labeled data, offering a detailed understanding of
scammer strategies, victim experiences, and responses. Below is an analysis based on the
percentages derived from the visualizations.

16
Figure 5: Combined visualizations: Scammer intent (pie chart), victim sentiment, impact, re-
porting status, and recovery type (bar charts).

1. Scammer Intent
Analysis: The pie chart categorizes scammer intent into four types: emotional (7.4%), finan-
cial (35.9%), both emotional and financial (26.5%), and unclear (30.3%). Financial scams
dominate, reflecting the primary goal of monetary gain. Emotional scams are less common but
often involve manipulation through relationships or trust-building tactics. The ”both” category
highlights scams that leverage emotional manipulation to achieve financial exploitation. The
”unclear” category suggests ambiguity in victim descriptions or insufficient details in posts.
Interpretation: Scammers predominantly target victims financially, often using emotional ma-
nipulation as a complementary strategy. The significant percentage of ”unclear” posts indicates
the need for clearer descriptions or better reporting mechanisms.

2. Victim Sentiment
Analysis: Victim sentiment is categorized into mixed emotions (dominant), sadness, helpless-
ness, anger, and regret. Mixed emotions account for the largest percentage (˜60%), reflecting
the complexity of victim experiences. Anger (˜15%) and sadness (˜10%) are also prominent
sentiments, indicating frustration and emotional distress.
Interpretation: Victims often experience a combination of emotions after being scammed,
highlighting the psychological impact of scams beyond financial losses. Anger and sadness
suggest feelings of betrayal and loss, while regret is less common, indicating victims may
blame external factors rather than themselves.

3. Level of Impact
Analysis: The level of impact is divided into high (˜2%), medium (˜80%), and low (˜18%).
Medium-impact scams dominate, suggesting that most scams cause moderate disruptions or
financial losses rather than catastrophic outcomes.

17
Interpretation: While most scams result in moderate losses, high-impact cases highlight se-
vere consequences for victims. This emphasizes the importance of early detection and preven-
tion to mitigate larger losses.

4. Reporting Status
Analysis: The reporting status is divided into ”reported” (˜33%) and ”not reported” (˜67%). A
majority of victims did not report scams, reflecting a lack of trust in authorities or unawareness
about reporting mechanisms.
Interpretation: The low reporting rate underscores the need for awareness campaigns to en-
courage victims to report scams and seek help. Improving trust in law enforcement and provid-
ing accessible reporting channels could increase these numbers.

5. Recovery Type
Analysis: Recovery efforts are categorized into no recovery (˜85%), therapy (˜5%), and finan-
cial recovery (˜10%). Most victims did not recover from their losses emotionally or financially.
Interpretation: The lack of recovery actions suggests gaps in support systems for scam vic-
tims. Therapy and financial recovery are rare, indicating limited access to resources or knowl-
edge about recovery options. This highlights the need for better victim assistance programs and
awareness about recovery mechanisms.

Interpretations from Visualizations


• Financial motives dominate scammer intent, with emotional manipulation often used as
a tool for exploitation.
• Victims experience a range of emotions post-scam, with mixed feelings being the most
common.
• Medium-impact scams are prevalent, while high-impact cases highlight severe conse-
quences.
• Low reporting rates emphasize the need for improved awareness and trust in reporting
mechanisms.
• Minimal recovery actions indicate gaps in support systems for scam victims.
These insights can inform strategies for scam prevention and victim assistance programs while
addressing gaps in reporting and recovery processes.

18
Step 4: Comprehensive Analysis of Scam Outcomes, Target
Groups, Remedies, and Victim Blame
This step consolidates the analysis of scam outcomes, targeted victim groups, scam types, reme-
dies initiated by victims, and victim blame attribution. The visualizations provide a detailed
understanding of scammer strategies and victim responses.

1. Scam Outcomes
Analysis: Scams were categorized as “Success” or “Failed.” A scam was labeled as “Success”
if the victim suffered financial loss or personal data theft, while “Failed” indicated that the scam
attempt was thwarted.
• 98.2% Success: Most scams achieved their intended goals.
• 1.8% Failed: Only a small fraction of scams were avoided by victims.
Interpretation: The high success rate highlights the effectiveness of scammers’ tactics and
underscores the need for stronger preventive measures to reduce victimization.

2. Targeted Victim Groups


Analysis: Victims were grouped into three categories based on demographic references in the
posts:
• Men (23.8%): Posts explicitly mentioning men as targets.
• Multiple Targets (70.9%): Posts targeting multiple demographics or general popula-
tions.
• Unclear (5.3%): Posts where the target group was not specified or ambiguous.
Interpretation: Scammers primarily target broad audiences or multiple demographics to max-
imize their reach. Men are specifically targeted in nearly a quarter of cases, while unclear posts
highlight gaps in reporting demographic details.

3. Type of Scam
Analysis: Scams were classified into four categories based on their nature:
• Technical (47.6%): Scams involving phishing, fake links, or technical impersonation.
• Mixed/Unclear (38.2%): Scams with overlapping characteristics or insufficient details
for classification.
• Emotional (7.9%): Scams leveraging emotional manipulation such as romance scams.
• Financial (6.2%): Direct financial frauds like credit card scams or fake investments.
Interpretation: Technical scams dominate due to the increasing sophistication of cyber-based
fraud methods. Emotional and financial scams, though less common, often rely on trust-
building and manipulation to exploit victims.

19
4. Remedies Initiated
Analysis: Posts were labeled as “Yes” if victims mentioned taking any remedial actions post-
scam (e.g., reporting to authorities, seeking financial recovery) and “No” if no remedies were
initiated.
• 98.2% No: Most victims did not take any remedial action after being scammed.
• 1.8% Yes: Only a small fraction of victims actively sought recovery or reported the
incident.
Interpretation: The lack of remedial actions suggests gaps in victim awareness about re-
sources or a lack of trust in recovery mechanisms. This highlights an urgent need for campaigns
to educate victims about reporting scams and seeking support.

5. Victim Blame Attribution


Analysis: Victims were categorized based on whether they blamed technology, themselves, or
external factors for falling prey to scams:
• Blamed Technology (9.7%): Victims attributed their experience to technological vul-
nerabilities (e.g., insecure platforms).
• Blamed Self (1.2%): Victims blamed themselves for being naive or careless.
• Blamed External Factors (89.1%): Most attributed their experience to external factors
like scammer tactics rather than personal or technological failure.
Interpretation: Most victims do not blame technology or themselves but rather focus on ex-
ternal factors such as the sophistication of scammers’ tactics. The low percentage of self-blame
suggests that victims generally do not see themselves as responsible for being scammed, which
could indicate a lack of awareness about preventive measures.

6. Age-Based Victimization Patterns (Bar Chart)


Analysis: The bar chart categorizes victims into three groups based on age-related references
found in the posts:
• General: Posts that do not specify an age group but target a broad audience.
• Elderly: Posts mentioning older adults as targets, often due to perceived vulnerabilities
such as lack of technical knowledge or trust in others.
• Student/Youth: Posts referencing younger individuals or students, typically targeted
through online platforms like social media and e-commerce.
Interpretation:
• General Population: The largest group targeted by scammers, reflecting the broad reach
of scam campaigns designed to appeal to a wide audience.
• Elderly: Represent a significant portion of targeted victims, likely due to their suscepti-
bility to phone call scams, investment frauds, and social engineering tactics.
• Students/Youth: A smaller but notable group targeted through scams like online shop-
ping fraud and cryptocurrency schemes, leveraging their frequent use of digital platforms.

20
(a) Scam Outcomes (b) Targeted Victim Groups

(c) Types of Scam (d) Remedies Initiated

(e) Victim Blame Attribution (f) Age-Based Victimization Patterns


21
Figure 6: Visual breakdown of scam outcomes, victim profiles, and post-scam responses from
Reddit discussions.

You might also like