DEEPFAKES: IMPACT, DETECTION &
SOLUTIONS
Presenter's Script — April 2026
Speak naturally. Bold phrases = emphasis points. Source citations in blue.
SLIDE 1 — TITLE SLIDE
Good [morning/afternoon], Professor. Today we're going to talk about deepfakes — not
as a future technology, but as something that is already causing serious real-world
harm right now.
Our presentation covers three things: the impact deepfakes are already having —
through case studies, the technology behind how they are made and detected, and
finally the solutions being proposed at the platform, policy, and user level.
We've also conducted a survey with 145 respondents to understand how aware
people around us actually are. Let's get into it.
SLIDE 2 — THE RISE OF DEEPFAKES
Let's start with the scale of the problem.
In 2023, there were around 500,000 deepfake files circulating online. By 2025, that
number is projected to cross 8 million — that's a 900% annual growth rate.
📌 Source: DeepStrike / Keepnet Labs — Deepfake Statistics 2025
Identity fraud attempts using deepfakes surged by 3,000% in a single year. And
deepfake incidents overall increased 10 times between 2022 and 2023.
📌 Source: Analytics Insight — Corporate Deepfake Fraud: The Rise of AI-Powered Financial
Scams (2025)
What makes this especially alarming is how easy it has become. You no longer need to
be a researcher or even a programmer. Free tools online can clone someone's voice
from just 3 seconds of audio, and swap faces in real time on a webcam. The barrier to
entry has effectively collapsed.
💡 Pause here to let the numbers sink in. The 3,000% stat usually gets a reaction.
SLIDE 3 — THREE DOMAINS OF HARM
Deepfakes cause harm in three main areas — financial, political, and social — and our
three case studies cover one each.
On the financial side: companies are losing millions. The average cost per deepfake
incident is around $500,000, and a single firm lost $25.6 million in one video call. We'll
cover that case in detail.
On the political side: a deepfake of Hillary Clinton was seen by thousands of people,
and 60% of viewers believed it was real — even though it was labeled as fake.
On the social side: 96% of all deepfakes online are non-consensual pornography,
and 99% of those target women. Taylor Swift's case showed just how fast this can
spread — 47 million views in 17 hours.
📌 Source: Kira, B. (2024) — Computer Law & Security Review; Momeni, M. (2025) — Journal of
Creative Communications
These aren't isolated incidents. They represent three very different ways deepfakes are
being used as a weapon.
SLIDE 4 — CASE STUDY 01 — ARUP ($25.6 MILLION)
This is one of the most striking corporate fraud cases in recent history.
In January 2024, a finance employee at Arup's Hong Kong office received what
seemed like a suspicious phishing email. He was rightly skeptical. But then — and this
is the key part — the fraudsters arranged a video conference call.
On that call, he saw and heard people he recognized: the CFO, and several senior
colleagues. They told him to authorize wire transfers. So he did. Fifteen transfers.
Totaling HK$200 million — roughly $25.6 million.
The fraud was only discovered when he followed up with Arup's actual UK headquarters
afterward. They had no idea any meeting had taken place.
📌 Source: CNN Business (2024); Fortune (2024); PRMIA Case Study Report (2025)
Here's what makes this case so important: no systems were hacked. No malware, no
stolen passwords. All of Arup's technical defenses were fully intact. The attack worked
entirely by manipulating human trust.
The fraudsters had used AI to generate real-time video and audio likenesses of Arup
executives — most likely from publicly available conference recordings and webinars.
📌 Source: Arup CIO Rob Greig described this as 'technology-enhanced social engineering' —
Adaptive Security (2024)
And here's the sobering stat: human accuracy in detecting high-quality video
deepfakes is only 24.5%. So expecting employees to spot a fake during a live call is
not realistic.
📌 Source: iProov (2025) — Deepfake Detection Accuracy Study
💡 Takeaway to emphasize: The problem is not the technology of your defenses — it's the trust
model those defenses rest on.
SLIDE 5 — CASE STUDY 02 — TAYLOR SWIFT (NON-CONSENSUAL DEEPFAKES)
In January 2024, AI-generated sexually explicit images of Taylor Swift were created
and shared across social media — primarily X, formerly Twitter, and 4chan.
One single post on X remained live for 17 hours and accumulated over 47 million
views, 24,000 reposts, and hundreds of thousands of likes — before the account was
finally suspended.
📌 Source: Kira, B. (2024) — Computer Law & Security Review, Vol. 54
X's response was to temporarily block all searches for Taylor Swift's name — which
is an extraordinary emergency measure that acknowledges how badly their content
moderation had failed.
Now, why is this a legal case study and not just a celebrity story? Because it exposed a
massive gap in the law. At the time, the United States had no federal law specifically
making deepfake intimate images illegal. A patchwork of state laws existed, but they
were inconsistent. Section 230 of the Communications Decency Act shielded the
platform from liability.
📌 Source: Kira (2024); Georgetown Journal of Gender and the Law (2024)
And the context is even darker: according to Kira's research, 96% of all deepfakes
online are non-consensual pornography, and 99% of those target women. The
Swift case happened to a famous person with a legal team and 200 million fans — for
ordinary women, the consequences are often irreversible.
📌 Source: Industry analysis of 14,678 deepfake videos — cited in Kira (2024)
💡 Emphasize: Swift's case triggered legislative action — multiple bills were introduced in Congress
within weeks. The TAKE IT DOWN Act was eventually signed in May 2025.
SLIDE 6 — CASE STUDY 03 — HILLARY CLINTON (POLITICAL DEEPFAKES)
On April 11, 2023, a 16-second deepfake video was posted on Instagram. In it, Hillary
Clinton — a prominent Democrat — appeared to endorse Republican Governor Ron
DeSantis for the 2024 presidential race.
The video was created using voice cloning and lip-sync technology. And here's the
critical detail: the post was captioned with the hashtag #deepfake. It was clearly
labeled. And yet...
Researchers analyzed approximately 1,670 comments on the post. 60% of those
commenters believed the video was real.
📌 Source: Momeni, M. (2025) — Artificial Intelligence and Political Deepfakes: Shaping Citizen
Perceptions Through Misinformation. Journal of Creative Communications, 20(1), 41–56
And the voting impact? Around 10% of those who believed it said explicitly that they
would no longer vote for DeSantis as a result — because they interpreted Clinton's
endorsement as a sign he was secretly a Democrat.
This ties into a really interesting psychological dynamic: Republican voters were more
susceptible to believing the video, not less — because it aligned with existing distrust
of DeSantis. Misinformation is most effective when it confirms what people already
suspect.
📌 Source: Momeni (2025), citing Nyhan & Reifler (2010) on confirmation bias
The video spread to X, YouTube, and TikTok before any platform moderation kicked in.
The case also references how Elon Musk shared a deepfake of Kamala Harris with just
a laughing emoji — that post got 130 million views.
The lesson: disclosure labels don't work. Putting #deepfake in a caption is not a
safeguard. The brain trusts what it sees and hears before it processes text.
💡 Good discussion point: Even Pew Research (2023) found that 50% of American adults aren't sure
what a deepfake even is.
SLIDE 12 — SURVEY INSIGHTS — KEY STATISTICS
We conducted a survey in April 2026 with 145 respondents — a 104% increase from
our previous wave of 71. The dominant age group was 18–24, making up 80% of
respondents, with the remainder from older groups.
📌 Source: Our own Deepfake Awareness Survey — April 4–20, 2026
The headline numbers are on screen, but let me walk through the most important ones.
88.3% of respondents said they are aware of what a deepfake is — up from 83.1%
in the previous wave. So awareness is growing.
66.9% said they have definitively seen a deepfake — up from 62%. Combine that
with the 14.5% who think they may have, and about 81% of our sample has likely been
exposed.
3.14 out of 5 is the average self-rated detection confidence. This is above the
midpoint — which sounds positive — but as we'll see on the next slide, this confidence
is dangerously misplaced.
55.2% said they trust no platform to moderate deepfakes — up from 47.9%. So not
only are people skeptical of content, they're skeptical of the systems meant to police it.
This is exactly the environment where misinformation thrives.
And 80.7% support mandatory labelling of all AI-generated content — a near-
unanimous result that has strengthened across both survey waves.
SLIDE 13 — THE OVERCONFIDENCE GAP & POLICY STANCE
This is the most concerning finding from our survey.
Respondents who said they 'know deepfakes well' averaged a detection confidence of
3.50 out of 5. Those who had never heard the term averaged just 1.82 out of 5. That's
a 92% gap — up from 66% in our previous wave.
📌 Source: Our Deepfake Awareness Survey — Section 4, Overconfidence Analysis (April 2026)
On the surface, you might think: great, the more aware people are, the more confident
they are. But here's the problem. Empirical research by Kobis et al. (2021) consistently
shows that even self-rated confident individuals perform near-chance on high-quality
deepfakes in controlled tests. So the confidence is not matched by actual detection
ability.
📌 Source: Kobis et al. (2021) — cited in our Deepfake Awareness Survey Analysis Report
In other words, the more someone thinks they can spot a deepfake, the more
dangerous they may actually be — because they're less likely to verify before acting or
sharing.
On the policy side: 80.7% of respondents support mandatory AI labelling without
exception. Only 2.1% said labelling is unnecessary. This signals strong public appetite
for regulatory action — which is relevant context for the solutions section coming up.
SLIDE 7 — SECTION DIVIDER — HOW DEEPFAKES ARE MADE
Now that we've seen what deepfakes can do, let's look at how they actually work —
because understanding the technology is important for understanding why detection is
so hard.
SLIDE 8 — GENERATION TECHNOLOGIES — GANS VS DIFFUSION MODELS
There are two main types of AI used to generate deepfakes. The first is Generative
Adversarial Networks, or GANs. The second, and more recent, is Diffusion Models.
GANs work through a kind of competition. You have two AI networks — a Generator
that creates fake images, and a Discriminator that tries to spot the fakes. They fight
each other until the Generator gets so good the Discriminator can't tell the difference
anymore.
📌 Source: DeepFake Generation Case Study PDF — Technical Analysis Part A
GANs are fast and widely used — tools like DeepFaceLab and FaceSwap are GAN-
based. But they leave detectable fingerprints in the frequency domain of images, which
makes them somewhat easier to detect.
Diffusion Models are newer and work differently. They take a real image, gradually
add random noise until it's completely scrambled, and then train an AI to reverse that
process — learning to 'undo' noise step by step. Tools like Stable Diffusion, DALL-E 3,
and Midjourney use this approach.
The key difference for detection: diffusion models produce far fewer statistical
artifacts. Most detection systems were designed for GAN-era content, and they
perform near-chance on diffusion outputs.
📌 Source: Technical Analysis Part B — Deepfake Detection: Methods, Limitations & The Arms
Race
Even more alarming: with diffusion models, you need just 20 reference images and
about an hour on a consumer-grade computer to build a personalized deepfake
generator of someone. No machine learning expertise required.
SLIDE 9 — FACE SWAPPING & VOICE CLONING PIPELINES
The slide shows two pipelines: one for face swapping, and one for voice cloning.
For face swapping, the process starts with a source video and a reference image of
the target person. AI encodes the identity of the target, then decodes it onto the source
person's facial movements and expressions. Tools like SimSwap and InSwapper can do
this from a single reference photo — in real time on a smartphone.
📌 Source: Technical Analysis Part A — Face Swapping section
For voice cloning, the AI captures the unique acoustic characteristics of someone's
voice — their tone, accent, and rhythm — and uses a text-to-speech model to make
them say anything. Models like ElevenLabs and XTTS can do this from as little as 3
seconds of audio with high fidelity.
📌 Source: McAfee (2024) AI Voice Scam Research Study — cited in Arup Case Study PDF
The threat that keeps security researchers up at night: combine these two in real time.
Tools like Deep Live Cam and RVC can now do live face swapping and voice cloning
together, with a total latency of under 400 milliseconds — below what humans can
perceive as a delay. This is exactly how the Arup fraud worked.
💡 This is a good moment to pause and let it sink in that someone could impersonate a person's face
and voice live on a video call, and it would be nearly undetectable in real time.
SLIDE 10 — DETECTION METHODS
Detection is the other side of the coin. There are four main approaches to detecting
deepfakes.
The first is CNN-based detection — using convolutional neural networks trained on
datasets like FaceForensics++. Models like EfficientNet-B4 can achieve 92–97%
accuracy on the dataset they were trained on. The problem? When you test them on a
different dataset or a different type of deepfake, accuracy drops to 55–65% — barely
better than flipping a coin.
📌 Source: Technical Analysis Part B — CNN-Based Detection section; FaceForensics++ (Rossler
et al., 2019)
The second approach is temporal analysis — looking at things like eye blink rate, head
movement patterns, and facial action units over time. Early deepfakes almost never
blinked because training data was biased toward open-eyed faces. Modern models
have improved, but subtle abnormalities in blink timing and muscle movement patterns
remain detectable.
The third is audio analysis — voice synthesis models leave fingerprints in the audio
spectrum that classifiers trained on real vs cloned audio can pick up, with 85–93%
accuracy. But heavily compressed audio destroys this signal, which is a problem on
social media.
The fourth is metadata forensics — checking whether an image has the expected
camera noise pattern, JPEG compression artifacts, or EXIF data. But metadata can be
stripped in seconds, so this is more of a supporting clue than a definitive method.
Combining all four methods in an ensemble gets you 94–98% accuracy — but only on
known generators. And the latency is 200–800 milliseconds per frame, which is far too
slow for real-time live call detection.
📌 Source: Technical Analysis Part B — Detection Method Comparison Table
SLIDE 11 — THE DETECTION-GENERATION ARMS RACE
The timeline on this slide tells an important story. Detection has always been playing
catch-up with generation.
In 2017, the first face-swap GANs appeared on Reddit. Detection was manual
inspection. By 2018–19, accessible tools like FaceSwap and DeepFaceLab were out,
and CNN-based detectors like XceptionNet were built in response. By 2021–22,
diffusion models and one-shot voice cloning were available — and detectors were still
catching up. By 2024–25, sub-second live face and voice swap on mobile phones
exists — and there is still no real-time detector that can reliably catch it.
📌 Source: Technical Analysis Part B — Arms Race Timeline
There's a structural reason this gap keeps growing: a new generator only needs to
fool the detector once to be useful. A detector needs to be right for every unknown
input, including deepfakes from generators it has never seen before.
Published detection models also suffer from a 12–18 month publication lag — by the
time a paper describes how to detect today's deepfakes, the state-of-the-art generators
have already moved on.
💡 Key insight: Detection alone cannot solve this problem. It's one layer in a multi-layer response —
not the answer.
SLIDE 14 — CODE IMPLEMENTATION
For the technical implementation part of our project, we built a deepfake detection
pipeline using open-source tools.
Our primary detector is EfficientNet-B4, pre-trained on FaceForensics++ — the
standard benchmark dataset with 1,000 videos manipulated using five different forgery
methods at three compression levels.
📌 Source: FaceForensics++ (Rossler et al., 2019) — Technical Analysis Part B
We trained on the c23 (heavily compressed) quality setting, which is closest to real-
world social media video conditions. We then evaluated the model cross-dataset on
DFD and Celeb-DF v2 — because in-domain accuracy alone is misleading. A model
that scores 99.7% AUC on its own test set can drop to 55–63% on a different dataset
— barely above random chance.
📌 Source: Technical Analysis Part B — Cross-Dataset Generalization Failure section
We also used OpenFace 2.0 for facial action unit extraction — detecting unnatural
muscle movement patterns over time — and Wav2Vec2 for audio artifact detection on
cloned voices.
For explainability, we implemented Grad-CAM visualizations — these show which
regions of the face the model is actually looking at when it makes a detection decision.
This makes the model's behavior interpretable and forensically meaningful.
And we performed a false positive analysis by demographic subgroup, because
existing detectors show systematic bias — higher false positive rates for darker skin
tones and East Asian faces due to underrepresentation in training data.
SLIDE 15 — SOLUTIONS FRAMEWORK — THREE PILLARS
No single solution can fix the deepfake problem. What's needed is a layered
architecture — and our proposed framework has three pillars.
Pillar 1 — Platform-level solutions. Platforms are the primary distribution channel for
deepfakes, so they must also be the first line of defense. This means mandatory
labelling of AI-generated content — both visible watermarks for users and embedded
metadata for machines. It means faster takedown windows: the US TAKE IT DOWN
Act (May 2025) requires platforms to remove flagged non-consensual intimate
deepfakes within 48 hours of notification.
📌 Source: Proposed Non-Technical Solutions PDF — Section B; US TAKE IT DOWN Act (May
2025)
It also means risk scoring — algorithmically assessing content for synthetic probability,
identity match, and harm potential before amplifying it. High-risk content should be
withheld from recommendation feeds pending human review.
Pillar 2 — Policy and legal solutions. The legislative landscape has changed
dramatically in the past two years. The EU AI Act requires transparent labelling with
fines up to 6% of global annual turnover. The DEFIANCE Act in the US establishes
civil damages of up to $150,000–$250,000 for victims. The Tennessee ELVIS Act
establishes a person's voice and likeness as protected personal property.
📌 Source: Proposed Non-Technical Solutions PDF — Section C; EU AI Act Article 50 (August
2025)
As of early 2026, 46–48 US states have some form of deepfake legislation, and
California alone has enacted 18 separate laws.
Pillar 3 — User education. Research from Huang & Hu (2025) found that technique-
based media literacy tips — specifically how to identify deepfakes — significantly
improved detection accuracy, and the effect lasted at least a week. Video format was
the most effective delivery method.
📌 Source: Proposed Non-Technical Solutions PDF — Section D, Media Literacy Programs; Huang
& Hu (2025)
The recommendation is to embed this in K-12 curricula, workplace cybersecurity
training, and community programs — particularly for older adults, who are the most
vulnerable demographic.
SLIDE 16 — FUTURE DIRECTIONS
Looking ahead, there are six areas we think deserve attention.
First, C2PA provenance authentication. This is a standard developed by Adobe,
Microsoft, and Sony that cryptographically signs media files at the point of creation.
Instead of asking 'prove this is fake,' it flips the burden: 'prove this is authentic.' This is
probably the most promising near-term technical solution.
📌 Source: Technical Analysis Part B — C2PA / Content Credentials section
Second, real-time detection research. The current gap is enormous — ensemble
detectors need 200–800ms per frame; live video requires under 33ms. Closing this gap
is a key research priority for the next few years.
Third, international coordination. Deepfakes don't respect borders. Political
deepfakes can be created in one country, hosted in another, and targeted at elections in
a third. Bilateral and multilateral agreements between democracies are needed.
📌 Source: Hillary Clinton Case Study — Regulatory Landscape section; Proposed Solutions PDF
Fourth, extending solutions to non-technical populations. Our survey showed that
detection confidence is already worryingly overestimated in a young, educated, digitally
active demographic. For older adults, rural populations, and less digitally fluent groups,
the risk is even higher. UNESCO has flagged this as a global education priority.
📌 Source: Proposed Non-Technical Solutions PDF — Section D.1, Awareness Campaigns; Our
Survey
Fifth, AI developer mandates. Apps like 'ClothOff' — which are specifically designed to
generate non-consensual intimate imagery — currently operate entirely outside any
regulatory framework. Platforms are regulated; the tools that create the content are not.
That gap needs to close.
Sixth, demographic-aware detection. Current detectors show significantly higher false
positive rates for darker skin tones and East Asian faces. This is a fairness issue that
needs to be addressed in training data and evaluation standards.
SLIDE 17 — CONCLUSION — THE WINDOW IS CLOSING
Let me bring this all together.
Deepfakes are not a future threat. They are a present one. A global engineering firm
lost $25.6 million in a single video call. A 16-second video of a political figure — clearly
labeled as fake — still convinced 60% of viewers and changed stated voting intentions.
Explicit AI-generated images of a celebrity reached 47 million people before a single
platform responded.
Our survey confirms that the population most likely to encounter deepfakes — young,
educated, digitally active users — overestimates their ability to detect them by a
92% margin. And 55% of them trust no platform to protect them.
📌 Source: Our Deepfake Awareness Survey (April 2026, n=145)
What this tells us is that the solution cannot rely on users spotting deepfakes, or on
platforms voluntarily acting, or on any single intervention. It needs to be layered:
technical provenance standards built into creation tools, platform obligations with real
enforcement, legal frameworks with meaningful penalties, and sustained public
education.
As the researchers put it — and we think this is the right framing — the window in which
democratic societies can establish effective countermeasures is finite and closing. The
goal is not to eliminate deepfakes. It's to build systems resilient enough to survive them.
📌 Source: Hillary Clinton Case Study — Conclusion section (Momeni, 2025)
Thank you. We're happy to take questions.