0% found this document useful (0 votes)
7 views28 pages

Dewan Ziaul Karim: Document Details

The document presents a thesis titled 'Seeing Beyond the Fake: Detecting Deepfakes Using Deep Learning-Based Computer Vision,' submitted for a B.Sc. in Computer Science at Brac University. It discusses the challenges posed by deepfake technologies and proposes a robust detection framework that combines temporal and visual analysis to enhance detection accuracy. The work aims to contribute to the development of reliable deepfake detection systems for applications in digital forensics and media verification.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views28 pages

Dewan Ziaul Karim: Document Details

The document presents a thesis titled 'Seeing Beyond the Fake: Detecting Deepfakes Using Deep Learning-Based Computer Vision,' submitted for a B.Sc. in Computer Science at Brac University. It discusses the challenges posed by deepfake technologies and proposes a robust detection framework that combines temporal and visual analysis to enhance detection accuracy. The work aims to contribute to the development of reliable deepfake detection systems for applications in digital forensics and media verification.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Page 1 of 28 - Cover Page Submission ID trn:oid:::1:3473048642

Dewan Ziaul Karim


DF
P2 Spring 2026

Conference Paper Check

CSE

Document Details

Submission ID

trn:oid:::1:3473048642 26 Pages

Submission Date 7,816 Words

Feb 4, 2026, 11:37 PM GMT+5:30


47,731 Characters

Download Date

Feb 4, 2026, 11:45 PM GMT+5:30

File Name

etecting_Deepfakes_Using_Deep_Learning_Based_Computer_Vision.pdf

File Size

202.4 KB

Page 1 of 28 - Cover Page Submission ID trn:oid:::1:3473048642


Page 2 of 28 - AI Writing Overview Submission ID trn:oid:::1:3473048642

41% detected as AI Caution: Review required.

The percentage indicates the combined amount of likely AI-generated text as It is essential to understand the limitations of AI detection before making decisions
well as likely AI-generated text that was also likely AI-paraphrased. about a student’s work. We encourage you to learn more about Turnitin’s AI detection
capabilities before using the tool.

Detection Groups
37 AI-generated only 41%
Likely AI-generated text from a large-language model.

0 AI-generated text that was AI-paraphrased 0%


Likely AI-generated text that was likely revised using an AI-paraphrase tool
or word spinner.

Disclaimer
Our AI writing assessment is designed to help educators identify text that might be prepared by a generative AI tool. Our AI writing assessment may not always be accurate (i.e., our AI models
may produce either false positive results or false negative results), so it should not be used as the sole basis for adverse actions against a student. It takes further scrutiny and human
judgment in conjunction with an organization's application of its specific academic policies to determine whether any academic misconduct has occurred.

Frequently Asked Questions

How should I interpret Turnitin's AI writing percentage and false positives?


The percentage shown in the AI writing report is the amount of qualifying text within the submission that Turnitin’s AI writing
detection model determines was either likely AI-generated text from a large-language model or likely AI-generated text that was
likely revised using an AI paraphrase tool or word spinner.

False positives (incorrectly flagging human-written text as AI-generated) are a possibility in AI models.

AI detection scores under 20%, which we do not surface in new reports, have a higher likelihood of false positives. To reduce the
likelihood of misinterpretation, no score or highlights are attributed and are indicated with an asterisk in the report (*%).

The AI writing percentage should not be the sole basis to determine whether misconduct has occurred. The reviewer/instructor
should use the percentage as a means to start a formative conversation with their student and/or use it to examine the submitted
assignment in accordance with their school's policies.

What does 'qualifying text' mean?


Our model only processes qualifying text in the form of long-form writing. Long-form writing means individual sentences contained in paragraphs that make up a
longer piece of written work, such as an essay, a dissertation, or an article, etc. Qualifying text that has been determined to be likely AI-generated will be
highlighted in cyan in the submission, and likely AI-generated and then likely AI-paraphrased will be highlighted purple.

Non-qualifying text, such as bullet points, annotated bibliographies, etc., will not be processed and can create disparity between the submission highlights and the
percentage shown.

Page 2 of 28 - AI Writing Overview Submission ID trn:oid:::1:3473048642


Page 3 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

Seeing Beyond the Fake: Detecting Deepfakes Using Deep


Learning-Based Computer Vision

by

Arobi Alamin Twinkle


22299058
Tanvir Bashar Aurpo
22201903
Sifat Alam
22299034
Syed Aiman Aabdee
22201462

A thesis submitted to the Department of Computer Science and Engineering


in partial fulfillment of the requirements for the degree of
[Link]. in Computer Science and Engineering

Department of Computer Science and Engineering


Brac University
Month Year.

© 2025. Brac University


All rights reserved.

Page 3 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642


Page 4 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

Declaration
It is hereby declared that

1. The thesis submitted is my/our own original work while completing degree at
Brac University.

2. The thesis does not contain material previously published or written by a


third party, except where this is appropriately cited through full and accurate
referencing.

3. The thesis does not contain material which has been accepted, or submitted,
for any other degree or diploma at a university or other institution.

4. We have acknowledged all main sources of help.

Student’s Full Name & Signature:

AROBI AL AMIN TWINKLE TANVIR BASHAR AURPO


22299058 22201903

SIFAT ALAM SYED AIMAN AABDEE


22299034 22201462

i
Page 4 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 5 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

Approval
The thesis/project titled “Seeing Beyond the Fake: Detecting Deepfakes Using Deep
Learning-Based Computer Vision” submitted by
1. Arobi Alamin Twinkle (22299058)
2. Tanvir Bashar Aurpo (22201903)
3. Sifat Alam (22299034)
4. Syed Aiman Aabdee (22201462)
of Summer, 2025 has been accepted as satisfactory in partial fulfillment of the re-
quirement for the degree of [Link]. in Computer Science in Summer 2025.

Examining Committee:

Supervisor:
(Member)

Dewan Ziaul Karim

Senior Lecturer
Department of Computer Science and Engineering
Brac University

Thesis Coordinator:
(Member)

Dr. Md. Golam Rabiul Alam

Professor
Department of Computer Science and Engineering
Brac University

Head of Department:
(Chair)

Dr. Sadia Hamid Kazi

Chairperson
Department of Computer Science and Engineering
Brac University

ii
Page 5 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 6 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

Abstract
In recent years, the proliferation of deepfake technologies has posed significant chal-
lenges to the integrity of digital media, public trust, and security. Deepfakes, gener-
ated using advanced deep learning models such as Generative Adversarial Networks
(GANs), can fabricate hyper-realistic audio-visual content, often making it indis-
tinguishable from genuine footage to the human eye. Deepfake detection focuses
on identifying media that has been altered, often using tools like generative adver-
sarial networks (GANs). While several detection methods have emerged, many are
limited by their reliance on spatial inconsistencies or specific artifacts that more
sophisticated generation techniques can easily mask. This paper proposes a robust
framework for deepfake detection that leverages both temporal and vision-based
analysis to identify subtle inconsistencies across video sequences. By incorporating
temporal dynamics—such as motion anomalies, blink rates, and micro-expression
irregularities—alongside frame-level visual cues like facial landmarks, texture incon-
sistencies, and illumination mismatches, our approach aims to enhance detection
accuracy against a wide range of deepfake types. The proposed approach prioritizes
controlled model capacity, explicit regularization, and fair comparative evaluation to
mitigate overfitting and dataset bias. By analyzing training behavior and architec-
tural sensitivity, this work establishes a robust foundation for subsequent extension
toward spatio-temporal and video-level deepfake detection. The findings contribute
toward developing more reliable and scalable deepfake detection systems suitable
for digital forensics, media verification, and online content moderation.

Keywords: Deepfake detection, Temporal analysis, Visual analysis, Computer vi-


sion, Image-based analysis, CNN, Fake media detection.

iii
Page 6 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 7 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

Table of Contents

Declaration i

Approval ii

Abstract iii

Table of Contents iv

1 Introduction 2
1.1 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2
1.2 Rational of the Study or Motivation . . . . . . . . . . . . . . . . . . . 2
1.3 Problem Statement . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.4 Objective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.5 Methodology in Brief . . . . . . . . . . . . . . . . . . . . . . . . . . 4
1.6 Scopes and Challenges . . . . . . . . . . . . . . . . . . . . . . . . . . 5

2 Literature Review 6
2.1 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
2.1.1 What Are Deepfakes? . . . . . . . . . . . . . . . . . . . . . . . 6
2.1.2 Deepfake Detection Process . . . . . . . . . . . . . . . . . . . 6
2.1.3 Challenges in Detection . . . . . . . . . . . . . . . . . . . . . 7
2.1.4 Ways to Improve Detection . . . . . . . . . . . . . . . . . . . 7
2.1.5 Datasets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
2.2 Review of Existing Research . . . . . . . . . . . . . . . . . . . . . . . 8
2.2.1 CNN-Based Approaches in Deepfake Detection . . . . . . . . . 8
2.2.2 Deepfake Detection Based on Hybrid Spatio-Temporal Models 9
2.2.3 Vision Transformers and Multimodal Approaches . . . . . . . 10
2.2.4 Datasets and Benchmarking . . . . . . . . . . . . . . . . . . . 11
2.2.5 Future directions and lookabouts . . . . . . . . . . . . . . . . 11
2.3 Summary of Key Findings . . . . . . . . . . . . . . . . . . . . . . . . 12

3 Proposed Methodology 14
3.1 Design Process or Methodology Overview . . . . . . . . . . . . . . . . 14
3.1.1 Design Objectives and Engineering Intent . . . . . . . . . . . 14
3.1.2 Dataset-Driven Design Rationale . . . . . . . . . . . . . . . . 14
3.1.3 Engineering Workflow and Design Stages . . . . . . . . . . . . 15
3.1.4 Design Constraints and Assumptions . . . . . . . . . . . . . . 15
3.2 Preliminary Design and Model Specification . . . . . . . . . . . . . . 16
3.2.1 Model Design Alternatives Review . . . . . . . . . . . . . . . 16

iv
Page 7 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 8 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

3.2.2 EfficientNetB0: Efficiency-Oriented Design . . . . . . . . . . . 16


3.2.3 ResNet50: Residual Learning with High Capacity . . . . . . . 16
3.2.4 MobileNetV2: Lightweight Architecture . . . . . . . . . . . . . 16
3.2.5 XceptionNet: Specialization in Spatial Artifacts . . . . . . . . 17
3.2.6 Custom CNN Architecture Proposal . . . . . . . . . . . . . . 17
3.2.7 Dataset Description . . . . . . . . . . . . . . . . . . . . . . . . 17
3.2.8 Model Comparison Strategy . . . . . . . . . . . . . . . . . . . 17
3.3 Trade-off Analysis and Architectural Justification . . . . . . . . . . . 17
3.4 Overfitting Behavior and Dataset Sensitivity Analysis . . . . . . . . . 18
3.5 Alternative Design Approaches . . . . . . . . . . . . . . . . . . . . . 18
3.6 Simulation Environment and Experimental Considerations . . . . . . 18
3.7 Scope Limitations of the Present Design . . . . . . . . . . . . . . . . 18

4 Conclusion 19

Bibliography 21

1
Page 8 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 9 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

Chapter 1

Introduction

1.1 Background
During the past ten years, the breakthrough in deep learning has transformed visual
media generation. It is now a possibility to produce photo realistic human faces,
speech, gestures, using techniques like Generative Adversarial Network (GANs) and
autoencoders. These man-made media, often referred to as deepfakes have influ-
enced the distinction between truth and fake to the farthest [9]. Although deep-
fakes were initially developed for entertainment and accessibility—such as dubbing,
filmmaking, or virtual avatars—their misuse has rapidly expanded into misinforma-
tion, fraud, and identity theft [7]. The evolution of deepfake generation models has
been accompanied by research on their detection. Early efforts focused on visual
forensics—identifying pixel-level anomalies, blending artifacts, and inconsistencies in
lighting [3]. With the rise of large benchmark datasets like FaceForensics++, Celeb-
DF, and the DeepFake Detection Challenge (DFDC) [5], more sophisticated models
emerged, particularly convolutional neural networks (CNNs), recurrent networks,
and hybrid CNN–LSTM architectures. Modern approaches also explore transformer-
based models [10], attention mechanisms, and multimodal cues that combine spatial
and temporal information [16]. Despite substantial progress, the reliability of deep-
fake detectors in uncontrolled, real-world conditions remains limited. Models often
fail when facing unseen manipulation methods, strong compression, or domain shifts
between datasets [14]. These limitations motivate continued exploration of robust,
adaptive, and explainable detection strategies that integrate both vision-based (spa-
tial) and temporal (motion-based) cues [2].

1.2 Rational of the Study or Motivation


The motivation for this research lies in the rising influence of manipulated media on
social and political stability. Deepfake videos can convincingly impersonate public
figures, generate false news, and erode trust in digital communication [12]. Because
human perception is easily deceived, technological solutions must be able to authen-
ticate visual content objectively and efficiently. Many existing detection systems rely
on analyzing single frames using CNNs such as Xception Net or MesoNet [6]. While
these can identify artifacts like blurred edges or inconsistent skin tones, they ignore
the temporal context of videos. Temporal inconsistencies—such as unnatural eye-

2
Page 9 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 10 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

blinking, head-movement irregularities, or desynchronized lip motion—are strong


indicators of manipulation [4]. Incorporating temporal reasoning using LSTM, 3D
CNN, or Vision Transformer architectures improves robustness [15]. From an engi-
neering perspective, the availability of GPU acceleration through CUDA and cuDNN
has made it possible to train deep spatio-temporal models more efficiently [13]. Also,
fake media ethical and legal issues contribute to escalating the need to design de-
tection systems that can be used to support journalism, forensics, and moderating
online platforms [11]. Thus, this thesis will provide a vision-temporal model that
will have a chance to design a hybrid that will be able to detect subtle artifacts of
manipulation, both in videos and images.

1.3 Problem Statement


Current deepfake detection techniques face multiple shortcomings. Frame-based
CNNs capture local inconsistencies but fail to detect manipulations that are tempo-
rally coherent [6]. Conversely, temporal detectors may overlook fine-grained spatial
distortions [4]. Many models also overfit to specific datasets, resulting in poor
generalization to unseen fakes or real-world conditions [14]. The rapid evolution
of generation models such as StyleGAN 3 and diffusion-based techniques further
widens this gap [10]. Therefore, the main research issue may be formulated in
the following manner: In which way is it possible to design a system based on
a deep-learning computer-vision capable of the appropriate ability to detect the
media produced by deepfakes, incorporating both visual (space) and motion-based
(temporal) gatherings and ensuring a high generalization of the system and its high
operation efficiency? Sub-questions include: What is the combination of CNNs and
time models (LSTM, 3D CNN, ViViT) to work with videos [15][1]? What is the
impact of diversity of datasets on the performance and transferability of models
[5][16]? Which computational optimizations can be used to train large scale, both
efficiently, using CUDA [13]? The answers to these will facilitate stronger, scalable
deep fake detectors that will eventually be armed to respond to manualizations in
the future.

1.4 Objective
The general aim of this thesis is to come up with an improved hybrid deep-learning
model that will be used to detect fake videos and images based on both time and
space. There are specific objectives, namely:
1. To examine spatially mismatched manipulated faces with convolutional archi-
tecture like ResNet or EfficientNet, or Xception Net [6][8].
2. To add time sequence modeling (e.g., LSTM, GRU, 3D CNN or Vision Trans-
former) to identify motional-based anomalies such as blinking and mouthing
[4][15].
3. To make effective use of each Puthrough GPU acceleration (CUDA, cuDNN)
[13].

3
Page 10 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 11 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

4. A way to compare the performance on the various datasets, such as FaceForen-


sics +, Celeb-DF, and DFDC [5][14].

5. To test the limitations and generalization capacities in the unseen manipula-


tions and levels of compression [10][16][2].

All these goals are meant to drive a balanced data intensive framework that is both
highly detected and cross-dataset robust.

1.5 Methodology in Brief


This study follows an experimental deep-learning methodology comprising data
preparation, model design, training, and evaluation.

1. Dataset Selection: Benchmark datasets such as FaceForensics++, Celeb-


DF v2, and DFDC are selected for their diversity in manipulation types and
compression settings [5]. These datasets include identity swaps, expression
manipulations, and partial edits [16].

2. Preprocessing and Feature Extraction: Each video is decomposed into


frames, with facial regions detected via MTCNN or Dlib [9]. Faces are aligned,
normalized, and resized to a consistent resolution. Optical-flow or key-frame
selection may be employed to reduce redundancy [7]. For images, preprocessing
includes contrast normalization and augmentation [12].

3. Model Architecture: The proposed hybrid architecture combines a CNN


backbone for spatial features with a temporal module (LSTM / 3D CNN
/ Temporal Transformer) for sequence modeling [4][15]. Attention modules
or depthwise separable convolutions may be integrated to enhance feature
localization [1].

4. Training and Optimization: Training will use GPU acceleration with CUDA
and cuDNN [13]. Techniques such as dynamic learning-rate scheduling, data
augmentation (e.g., Face-Cutout [2]), and transfer learning will be applied to
improve performance and prevent overfitting.

5. Evaluation: The model will be assessed using metrics including accuracy,


precision, recall, F1-score, and AUC. Cross-dataset validation will measure
generalization [14]. Visualization tools like Grad-CAM will interpret model
focus regions [8].

Such a methodology will guarantee a systematic way of data collection through the
interpretation of models to make the results reproducible and allow a critical anal-
ysis.

4
Page 11 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 12 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

1.6 Scopes and Challenges


Scope

1. Dual-Modality Detection: The study targets both deepfake videos and


images, focusing primarily on facial manipulations such as identity swaps and
expression alterations [3][5].

2. Hybrid Spatio-Temporal Model: By combining CNNs for visual artifacts


and temporal networks for motion inconsistencies, the framework aims to cap-
ture both static and dynamic cues [4][15].

3. Benchmark Evaluation: Experiments will be conducted using standard


datasets: FaceForensics++, Celeb-DF, DFDC, and SID-Set—for comprehen-
sive validation [5][16][12].

4. GPU-Accelerated Implementation: CUDA and cuDNN support will al-


low scalable training and faster inference suitable for real-time or near-real-
time detection [13].

5. Ethical and Forensic Applications: The system’s potential extends to dig-


ital forensics, media authentication, and social-media moderation, promoting
accountability in AI-generated content [11].

Challenges

1. High-Quality Deepfakes: Modern generative models produce nearly artifact-


free fakes, reducing the effectiveness of purely visual detectors [10][16]. Dataset
Bias and Generalization: Training datasets often reuse the same identities,
causing overfitting and weak performance on unseen faces [2].

2. Computational Complexity: Spatio-temporal models demand significant


GPU memory and training time [13].

3. Cross-Domain Robustness: Detectors may fail on videos with new manip-


ulation styles, lighting, or compression artifacts [14].

4. Explainability and Transparency: Deep networks act as black boxes, lim-


iting interpretability. Visual explanation tools like Grad-CAM help but remain
imperfect [8].

5. Ethical and Privacy Concerns: Handling datasets with real identities re-
quires compliance with privacy and responsible-AI guidelines [11].

6. Adversarial Resistance: Attackers can intentionally craft adversarial per-


turbations to deceive detectors, necessitating ongoing adaptation [10].

By acknowledging these scopes and challenges, this thesis sets the foundation for
a balanced, ethical, and technically feasible exploration of deepfake detection using
deep-learning-based computer vision.

5
Page 12 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 13 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

Chapter 2

Literature Review

2.1 Preliminaries
2.1.1 What Are Deepfakes?
Deepfakes are computer-generated videos or images developed with sophisticated
artificial intelligence (AI) devices such as generative adversarial networks (GANs) or
autoencoders [3], [5], [14], [11]. Widely used programs like DeepFaceLab, FaceSwap,
or even StyleGAN are capable of swapping faces, altering facial expressions, or even
creating completely new ones that are highly realistic [3], [5], [14]. Such tools have
evolved to the point that they are highly difficult for naked-eyes to see a deepfake,
forcing researchers to devise methods to better detect them [5], [14], [11].

2.1.2 Deepfake Detection Process


Deepfake detection is a technology that enables detection of any evidence that sug-
gests a video or image has been manipulated using the methods of computer vision
and deep learning [5], [14], [11]. These hints may be either visual (such as unusual
textures on one picture) or motion-related (such as unnatural blinking on a video)
[3], [5]. Common approaches include: Image Visual Detection: Convolutional
Neural Networks (CNNs), including Xception or EfficientNet assess irregularities in
an image, such as the edges of the image not being straight or the lighting being
imperfect. Or, as an example, the Error Level Analysis (ELA) algorithm detects
the differences introduced by compression, at up to 88.2 percent accuracy when
combined with simple classifiers such as KNN [5], [14], [11], [1], [8]. Motion Clue
Detection: video model detectors, such as Long Short-Term Memory (LSTM) net-
works or 3D CNNs, work with video sequences and identify any unusual motion, such
as blinking irregularly or lipsync latency. Visual and motion cues combined have
95 percent accuracy when CNNs are used in conjunction with LSTMs [3], [5], [4],
[7]. High Vision Models Vision Transformer (ViTs) can identify items, such as eyes
or mouths, within images and videos, with 89.9 perent and 87.2 percent accuracy
respectively, by observing both the appearance and the motion pattern [10], [15].
Acceleration using GPUs: CNNs and other models can be trained much faster
using tools such as CUDA and cuDNN, reaching speeds under 1521 milliseconds to
process many images simultaneously, allowing them to be trained on large datasets
such as DFDC (470 GB) and spotted, or to identify an object in a single frame [7],

6
Page 13 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 14 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

[3], [14], [15], [11]. Although these techniques can achieve 98 percent accuracy in
laboratory work, they tend to fail on actual video, particularly where novel tricks of
deepfaking or compression come into play [5], [14], [11].

2.1.3 Challenges in Detection


There are a few issues in studying indicate that deepfake detection is a hard goal:
Low Flexibility: NeoNet models trained on datasets such as FaceForensics++
achieve 95-99 percent accuracy, but crater to 65-82 percent in tests on real-world
conditions, such as DFDC due to video style differences or recent advances in deep-
fake technology [5], [14], [11]. Condensed Videos: It is common in social media
where videos become poor in quality, concealing indicators which are used by de-
tectors. As an example, in one of the studies, the accuracy dropped to 92.3 on
compressed videos as opposed to an initial percentage of 98.5 that was on clear
videos [7]. Trick Attacks: Trick-attacks are formatted to bypass detectors, and
a lot of systems are not trained on trick-attacks [14], [2], [11]. Data Overlap: At
random split, data can be given out so that models will focus on memorizing faces
rather than learning true clues, and this will inflate results. This is one way that
has reduced errors by 35 percent [5], [14], [11]. Narrow Focus: The vast majority
of detectors merely scan faces, which overlooks fakes in voices, backgrounds or body
language [3], [14].

2.1.4 Ways to Improve Detection


To improve the detection researchers have proposed some solutions: Smart Data
Tweaks Face-Cutout: Face-Cutout creates a technique that withheld non- MLU
parts of an image during training so that a model specializes in making changes and
reduces errors by 35 percent or more when training on datasets like the Celeb-DF
dataset [1]. Attention: Adding instruments that have been built to accentuate
particular areas of the face, such as the eyes or the mouths, enhance detection. It
got 87.2 percent up to 92.5 percent balance score on video in a single study using this
to reach 87.2 percent [15], [1]. Display transparent explanations: Grad-CAM
diagrams enable visualizing of what aspects of a video or an image are being verified
by the model, which is more reliable when used as evidence in a court [7], [2], [1].
Combining Clues: combining visual, motion and sound clues, the system achieved
96.4 percent accuracy in certain datasets but showed poor performance when solving
other datasets (52 percent on NeuralTextures) [3], [16], [15]. Lightweight Models:
Simple, based on ELA and a naive classifier has achieved 88.2 percent accuracy
where power constraints are an issue. This kind of thing is used in devices with low
power available [8]. Fighting Trick Attacks: Training against trick attacks, such
as looking at frequencies, are good but even more research is needed [2], [11], [1].
These concepts assist, and it is still a challenge to make detectors fast, flexible, and
reliable in regards to all types of deepfakes. Test Data/ Metrics. Certain datasets
and measurements are taken by researchers to validate the effectiveness of detection
systems.

7
Page 14 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 15 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

2.1.5 Datasets
FaceForensics++: Has on the order of 1000 real and 4000 fake videos (Face2Face,
FaceSwaps), but the fakes are not as well-trained or advanced as those [5], [11].
DeepFake Detection Challenge (DFDC): A huge dataset of 470 GB with di-
verse videos, though having the capability of overlapping the face which will distort
the result [14], [11].
Celeb-DF: Has better than older datasets, (around) 5,639 realistic fake videos, but
not as natural [5], [14].
SID-Set: An image set (300,000 images, half real, half AI generated and edited) of
social media fakes but only includes images [16].
Others: Smaller datasets such as Deepfake-TIMIT or Yonsei are Real/Fake Face
sets that are not flexible because of their size or quality [3], [8]. Metrics
Accuracy: Indicates the rate of the model being able to recognize true or false
media which during test = 87-98 percent but during application it is lower [7], [3],
[10], [14], [4], [15], [11], [8].
Precision and Recall: To what extent does the model detect fakes without con-
fusing real ones (e.g. thinks that a good model is at least 90 percent).
F1-Score: Measures both accuracy and recall and best results of 92.5 percent on
Celeb-DF.
ROC-AUC: ‘Checks on variation by threshold, to 94 percent to 96 percent managed
systems.
Error Rate: Low error rates (11.14 percent): models with low error rates are
considered reliable.
Inference Time: Faster speed, optimized models can be less than 15 to 21 ms to
use in realtime [7], [3], [10], [14], [4], [15], [11], [8]. These measures and datasets are
useful to test out detection systems, though problems such as dataset diversity, and
practical performance are challenges.

2.2 Review of Existing Research


2.2.1 CNN-Based Approaches in Deepfake Detection
CNNs used to identify deepfakes have shown to be very effective in identifying the
spatial limits of modified content particularly in a frame-period by frame detection
method with attention placed on visual discontinuities such as differences in texture
and blending mistakes [5], [14], [11]. The use of CNNs based on the Xception archi-
tecture and trained on locating deepfish in videos in social media invented by Mitra
et al. assumes a sequence of actions that include the extraction of key frames but
analyzes the artifacts, including processing the variations in textures and a blend
misalignment, in a providing of compressed social media pieces in the real world.
This model trained on FaceForensics++ and on DFDC achieved a 98.5 per cent
and 92.3 per cent accuracy on uncompressed and highly compressed videos respec-
tively which observes the challenges of delivering service using a real-world model
where quality decisions are often fuzzy or poor-quality due to platform recordings,
and also outcompetes the traditional methods because it concentrates solely on the
portions of the face but not on aspects of temporal motion hence it is not as use-
ful as the more sophisticated deepfakes [7], [6]. This is more effective than the old

8
Page 15 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 16 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

techniques that use the areas of the face as the only factor and do not consider
the contribution of the temporal movement, which thwarts their application in the
absence of the superior deepfakes because of the capacity to retain the frames [7],
[6]. A comparable investigation by Nawaz et al. achieved improved outcomes and
decreased the reliance on the use of elaborate calculations on the basis of the Error
Level Analysis (ELA) alongside lightweight CNNs Basho CNNs and conventional
classifiers like KNN with stress on the area of transfer learning of feature extraction
by already processed 255x255 images. The analysis considered ELA to comprehend
compression differences of fake faces suggesting compare the capacity of CNN to
determine deep elements considering restricted resources, and when operated in the
Yonsei database (2041 images, 1081 genuine and 960 faked) it reached 88.2 percent
capacity, outperforming baselines (VGG16 88.46 and ResNet152 76.79), and giving
it appropriate to edge devices yet where the video sequences or modifying manip-
ulations are taken into corporation there exist deficiencies whereby the symptoms
The system mitigates instantaneously the computational effectiveness of the edge
devices but wherein of pursuing evolving manipulations or with treachery in view of
the video sequences there are disadvantages that the symptoms are scattered over
time [8].
An efficient prompt of scalable deepfake in cybersecurity: working memory An exam-
ple of an efficient CNN architecture is the EfficientNet, which represents an efficient
image detector, whose models such as EfficientNetV2L together with DNN classifiers
exhibit stable learning with minimal overfitting. Here, preprocessing, scaling a pic-
ture, normalizing, augmenting assists right the behaviour throughout trustworthy
working while employing the method of hyperparameter enhancement using dense
neural network classifiers, which received an exactness of about 92 percent, precision,
recall and F1-score on a stacked sample of 140,000 images which enables it to prevail
over DenseNet 201 (around 50 percent due to overfitting) and overperform with the
execution of traditional CNNs. However, such studies widely highlighted problems
in computing that arose with large amounts of data, or videos executable with time
to train volatile data, such as high demands on the amount of GPU resources in
volumetric data training. The research offered by surveys gives a comprehensive
analysis and assessment of the deficiency in efficiency of CNN. It is said that when
utilizing the mechanism of the handcraft, it poses notable difficulties in the high
resolution images when working in different environments [1], [5], [14], [11].

2.2.2 Deepfake Detection Based on Hybrid Spatio-Temporal


Models
This type of interaction between CNN and temporal network with LSTM models
as hybrids and LSTM models used simultaneously to detect stat and motion in-
consistency in deepfaking video is strongly promising with a significant possibility
of getting both a stat and motion anomaly challenged by frame-only analysis [3],
[4]. One of the earliest structured uses that used CNN-LSTM hybrids on deep
fake issues were made in the Nguyen et al. paper, which applied the hybrids to
a frame sequence analysis problem within datasets assembled by initiatives such as
FaceForensics++ and Deepfake-TIMIT, and added capsule networks and frequency-
domain features to render the hybrids robust. An approach that combined CNN
with visual corruption (e.g. distorted edges, strange reflections) and LSTM/GRU

9
Page 16 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 17 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

with time modularity (e.g. patterns of blinking) was also documented in this pa-
per, performing notably markedly whenever fakes are good in single but poor in
motion frames. The model is constructed by considering spatial coding strategies
that reduce the impact of compression and introduce capsule networks that enhance
the amount of generalization [3], [4]. Furthermore, recent advances on CNN-LSTM
designs have offered the prospect of video fragmentation and anomaly-loci with deep-
fakes, especially using optimization protocols such as MultiStepLR. The Rahman et
al system demonstrates the application of hybrid design configurations to address
the tough detection tasks whether to counterfeit motions or lip-synchronization by
MultiStepLR minimization and Grad-CAM visualization of findings, which are both
95 percent accurate on FaceForensics + and DFDC, the heatmap showing that com-
putation is paid on the regions of the images being edited rather than memorization,
and the heatmap is in comparison with other models using deepsight [12] of interest.
Footnote/foreach FC-DBNPL by Ram et al., combines Fuzzy C-Means mode of clus-
tering with Deep Belief Networks and pairwise learning (siamese architecture with
contrastive loss), Gabor filters and features (MSE, PSNR, SSIM, etc.), to achieve
an accuracy of 96.26-98.1 on FaceForensics++, FaceSwap, and DFDC, with normal
computation times (15.24-21.3 ms) lower than the competitors. Such cross-modes
congruency and matches demonstrate the impact of hybrids to fine-tune pattern
recognition that is information converting on the aspects of space and time [9].

2.2.3 Vision Transformers and Multimodal Approaches


As mentioned above, the limitations and drawbacks of CNN-based models have
given way to transformer-based and multimodal models, with modulated attention
mechanisms that handle the form of images in patches to enhance their accuracy
when focusing on inconsistencies [10], [16], [15], [11]. Ghita et al. tested Vision
Transformers (ViTs) to detect deepfakes in pictures based upon a combination of
self-attention and patch processing on a set of 40,000 images on Kaggle with a set
of suggested parameters such as a learning rate of 0.001 and a batch size of 256
running over 200 epochs. It tested inconsistencies in patches on an enhanced face
(89.91) and correct (89.91) with increasing preciosity and reliance in data efficiency
concerns archetypal of the deep fake-AI, with, however, transfer learning and fine-
tuning (learning rate 0.001, batch size 256, 200 moments), yet with a bias towards
detecting deepfakes on genuine pictures [10]. A proposed generic (facial landmarks,
DSC, CBAM only) ViViT achieved an accuracy and F1-score of 87.18 and 92.53
on Celeb-DF v2 (8, 000 videos) with numerous eyes, nose, and mouth targets and
few noise constraints, outperforming baselines such as MesoNet and Xception with
ablation, [15]. Multi-modal fusion This use of multi-modal fusion is increasingly
popularized as shown in Al Amin et al. where he is using audio style (MFCC) cues
when combining visual (CNNs), temporal (RNNs), and audio cues together to form a
complete cue using XceptionNet on the spatial sequences and RNNs on the temporal
sequences including associated biometric data. The model demonstrates that we can
use the different modalities to boost the recognition capacity by interpolating the
biometric and audio information which achieved the ratings of 96.36 percent on
Deepfakes and about 52 percent on NeuralTextures, beginning with improvements
such as attention-based operations and graph neural networks [3], [16]. Huang et al.
SIDA presents mask prediction and text explanations on SID-Set (300,000 images),

10
Page 17 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 18 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

which outperforms the baselines known to place edits or make large-scale predictions
on both social image localization and generalization, and classifies on large vision
language models and detects on unique tokens [16].

2.2.4 Datasets and Benchmarking


FaceForensics++ is the de-facto deepfake detection research benchmark data set.
It contains about 1-thousand valid videos and 4-thousand manipulative videos that
utilize such techniques as Face2Face and FaceSwap, annotated in case of spatial
and time anomalies, and compression levels and various manipulations of face such
Face2Face (205 videos), and FaceSwap (210 videos) [5], [11]. One other set of spe-
cialised datasets with either smaller scopes of deepfake generation include the DFDC
dataset (exceeding 470 Gb of most forms of videos of sizes, gender, race and condi-
tions) and the Celeb-DF set (exceeding 5639 high-quality fakes featuring celebrity
faces), and re-encoded subsets to test compression (c=15, 23, 40) [14], [5]. SID-Set
handles 300,000 samples (equipped with real, AI-generated, and edited images) in
the social media, whereas smaller ones gather Deepfake-TIMIT (620 fakes) and Yon-
sei (2,041 images), once again based on audio-visual or face-based tests, including
CelebA-HQ with still images [16], [3], [8]. This innovation has enabled the introduc-
tion of new opportunities in terms of deepfake monitoring with surveys suggesting
that initial datasets such as UADFV and DARPA MediFor are small and of poor
quality with the third generation dataset like DeeperForensics-1.0 being more realis-
tic and bias. At least some authors discussing the need to use benchmark challenges
like the DFDC competition in real time evaluations have found that there exists the
phenomenon of cross-dataset assessment but that performance suffers when trained
on one style and tested on another. However, the issue of generalization is extremely
large in such applications due to onboard biases in datasets and half-hearted com-
pression in datasets whereby limited view of unknown manipulations and adversarial
engineered deepfakes are covered. The need of complex data preprocessing and iden-
tity particle partitioning, low memory usage and conformance to the demands of the
real world is a vital disadvantage in the deepening of such structures, and the use
of MTCNN to extract visions of the face and OpenCV to acquire features. Sur-
veys entangle themselves recording that primitive collections have 95-99 well-being
in the framework in real-world tests the form 65-82 when hybridized planning and
opponent protections begin and that larger, richer, and more operateable datasets
as well as understandable mechanisms are as necessary. Accuracy, precision/recall
(less than 91 percent), F1-score, ROC-AUC (94 -96 percent), inference time (15-21
ms) could be used as measurements to benchmark, but issues like leakage of data
and absence of diversity is still there, no extensive cross-dataset explorations done
[5], [14], [11], [7], [3], [10], [4], [15], [8].

2.2.5 Future directions and lookabouts


Among the latest methods in detecting deepfakes, the main vulnerabilities in achiev-
ing these capabilities are being mitigated by potentially analyzing new offerings such
as frequency-domain analysis, adversarial training, and blockchain-based authentic-
ity verification to build robustness and generalization [5], [14], [11]. The spectral
inconsistencies overlooked by spatial models are identified with frequency-domain

11
Page 18 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 19 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

techniques like Discrete Cosine Transform (DCT) coefficients or a high-frequency


noise and create methods using ideas like CLRNet detect its presence with 93.8
percent precision on WildDeepfake [11]. Training on adversarial examples (i.e., on
crafted deepfakes e.g. a FGSM, cw-l2 attack attack) has been shown to enhance
resilience, where researchers report ratios up to 10 percent improvements in robust-
ness to unseen manipulations in adversarial training [14], [11]. According to surveys,
blockchain-based solutions provide the management of media provenance through
decentralized ledgers that provide a non-visual aspect of authenticity detection but
one that is yet to be practical to implement in practice [14].
Graph neural networks (GNNs) are directions of future research based on the results
from Al Amin et al. to determine how relationships across frames reflect moves on
facial features can be used to detect sophisticated manipulations such as full-body
manipulation or audio-based deepfakes [16]. Federated learningPrivacy-saving meth-
ods are being developed to allow targeted training without a breach of user data to
work across the board regarding the ethical issue in the practice [16]. Triplet train-
ing or multitask learning triplet training promote interpretable models and through
a trade off between accuracy and interpretability and variably reach Australian frac-
tion near 94 percent on DeeperForensics-1.0 with triplet training [11]. The methods,
however, have weaknesses on computational complexity and on working with more
and more diverse datasets to include non-facial manipulations and other real-world
perturbation e.g. noise and compression. The absence of standardisers merging
them with current spatial-temporal models also shows that there is a significant re-
search factor gap needed in scalable, vigorous, and interpretable deepfake detection
systems [5], [14], [11].

2.3 Summary of Key Findings


As can be seen, current deepfake detection experiments positively indicate that a
combination of visual (spatial) and motion-based (temporal) analysis beats mod-
els that utilize either of the mentioned types of features, as they are more both
consistent in detecting the existence of manipulations as well as those related to
open image artifacts in each case [7], [3], [14], [4], [15], [11]. FC-DBNPL (Fuzzy
C-Means clustering/pairwise learning), CNN-LSTM hybrid and ViViT (augmented
with attention) models have achieved an average accuracy of 87-98 percent on test
datasets including FaceForensics++ (meaning: 5,000 videos), DFDC (470 GB of
mixed content), and Celeb-DF (5,639 fake with high realism). As an example, FC-
DBNPL achieves 96.26 to 98.1 cues on spatial features (e.g., distorting texture with
MSE, PSNR, and SSIM) with 96.26-98.1 percent accuracy and less than 1.02-1.14
error rates, CNN-LSTM models best time-varying feature (e.g., blinking pixels nat-
uralness) with 95-98 percent accuracy, and ViViT concentrates in facial features
landmarks with 87.18 [9], [3], [4], [10], [15]. Yet, these models tend to fail when
applied in real life with recognition error dropping to 65-82 percent on natural data
with compression artifacts and the manipulation types not previously seen since
surveys obtain restrictions in generalization and defenses to adversarial attacks. A
possible solution to this involves generalized methods such as Face-Cutout (reduc-
tion in errors up to 35 percent by dynamically masking non-fake regions) to also
focus on manipulated parts and explainable methods such as Grad-CAM (display-
ing heatmaps of detected anomalies) to ensure greater transparency to forecasts.

12
Page 19 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 20 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

In general, spatial-temporal cue-based integration provides a firm detection frame


but the next-generation systems should be reinforced to accommodate compression,
different manipulations and computation requirements to extend their application
to the real world [5], [14], [11], [1], [2].

13
Page 20 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 21 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

Chapter 3

Proposed Methodology

3.1 Design Process or Methodology Overview


3.1.1 Design Objectives and Engineering Intent
This stage has a design process that is based on clear objectives established following
the problem statement and research gaps that were identified during Phase 1. As
opposed to the methods that are more attuned to the methods of optimization of the
performance, the study follows an engineering-based approach, which means that
both the theoretical and practical limitations will guide the architectural decisions
[5], [11], [14].
The main aim of the phase is that it will design, implement and test numerous
convolutional neural network (CNN) architectures to detect deepfakes, as well as
pretrained and custom-designed models. The research objective is to examine the
effects of the various architectural philosophies on detecting performance which in-
clude the deep residual learning, lightweight mobile oriented design and depthwise
separable convolution. Moreover, this step also analyses the effect of dataset fea-
tures on model behavior, especially against overfitting, bias, and cross-dataset per-
formance, as demonstrated in previous works on ResNet, MobileNet and Xception-
based detectors [6], [8], [10].
Another aim is the creation of a tailored CNN structure that is particularly aimed at
preventing overfitting by studying the model capacity, explicit regularization, and
attention procedures, which have been identified as major challenges in deepfake
detection research [5], [11], [14]. The correctness of all the proposed designs is tested
with simulation-based experiments which guarantee the functional correctness and
fair comparability. Together, these goals will help make the tasks carried out during
this step to be based on the principles of academic rigor and consider the engineering
limitations confronting reality [12], [14].

3.1.2 Dataset-Driven Design Rationale


Deepfake detection performance is highly dependent on the characteristics of the
datasets used for training and evaluation. Unlike conventional image classification
tasks, deepfake datasets vary significantly in manipulation quality, compression ar-
tifacts, visual realism, and identity overlap, which has been consistently reported as
a key limitation in detector generalization [5], [11], [14].

14
Page 21 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 22 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

To address this challenge, multiple datasets with distinct properties are deliberately
selected. The 140K Real and Fake Faces dataset provides large-scale diversity in
identities and visual conditions, allowing evaluation of scalability and training sta-
bility. In contrast, datasets such as HardFake vs Real Faces contain highly realistic
manipulations that test a model’s sensitivity to fine-grained artifacts. Additionally,
FaceForensics++ (C23) introduces compression artifacts representative of real-world
social media content.
By incorporating datasets with diverse characteristics, the design process emphasizes
cross-dataset robustness rather than dataset-specific optimization. This ensures that
experimental conclusions reflect architectural behavior rather than overfitting to a
particular dataset.

3.1.3 Engineering Workflow and Design Stages


The engineering workflow adopted in this study follows a structured and modu-
lar approach to ensure reproducibility, transparency, and fairness in comparative
analysis, consistent with established experimental practices in deepfake detection
research [11], [12]. The process begins with dataset exploration, where class bal-
ance, visual quality, and manipulation complexity are examined, as recommended
in benchmark-based studies [5], [14].
Next, all images are preprocessed and standardized. Inputs are resized to a fixed
resolution of 224×224 pixels and normalized to ensure compatibility with pretrained
CNN architectures and to eliminate input-related bias. Pretrained CNN models are
then introduced using transfer learning, serving as baseline engineering solutions to
evaluate established architectures within the deepfake detection context.
In parallel, a custom CNN architecture is designed with explicit emphasis on reduc-
ing overfitting and improving generalization. All models are trained under identical
experimental conditions within a simulation-based environment. Finally, perfor-
mance evaluation and comparative analysis are conducted using standardized met-
rics and qualitative observations of training behavior, ensuring that all designs are
assessed on equal terms.

3.1.4 Design Constraints and Assumptions


The design and experimentation in this phase are conducted under several practical
constraints. Computational resources are limited to cloud-based GPU environments,
imposing restrictions on training time, memory usage, and batch size, a limitation
widely acknowledged in deepfake detection experiments involving deep CNNs [11],
[13]. Additionally, public deepfake datasets often contain noise, identity overlap,
and labeling imperfections, which may influence model behavior and contribute to
overfitting [5], [14].
Architectural constraints also play a role, as excessively deep models tend to in-
crease computational cost and susceptibility to overfitting. A key assumption of
this phase is that frame-level image classification provides a reasonable approxi-
mation for analyzing deepfake detector behavior before extending the research to
video-level temporal modeling. This assumption allows the study to focus on spatial
feature learning, deferring temporal dependency analysis to Phase 3.

15
Page 22 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 23 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

3.2 Preliminary Design and Model Specification


3.2.1 Model Design Alternatives Review
This phase evaluates multiple CNN architectures rather than relying on a single so-
lution. Each selected model represents a distinct architectural philosophy, enabling
comprehensive engineering analysis. The models considered include EfficientNetB0,
ResNet50, MobileNetV2, XceptionNet, and a custom CNN architecture proposed in
this study. Together, these models cover a wide range of complexity, capacity, and
design intent [6], [8], [10].

3.2.2 EfficientNetB0: Efficiency-Oriented Design


EfficientNetB0 is based on a compound scaling strategy that balances network depth,
width, and input resolution [10]. From an engineering perspective, it represents
a carefully optimized trade-off between representational power and computational
efficiency. The architecture employs mobile inverted bottleneck convolution blocks
and squeeze-and-excitation mechanisms, enabling effective feature extraction with
relatively few parameters.
Empirical observations indicate that EfficientNetB0 demonstrates stable training be-
havior, particularly on large-scale datasets. However, its limited depth may restrict
its ability to capture extremely subtle artifacts present in highly realistic deepfakes.
Despite this limitation, EfficientNetB0 serves as a valuable efficiency-oriented base-
line.

3.2.3 ResNet50: Residual Learning with High Capacity


ResNet50 utilizes residual connections that facilitate deep hierarchical feature learn-
ing while mitigating vanishing gradient issues. Its depth enables the detection of
fine-grained texture inconsistencies and blending artifacts commonly found in ma-
nipulated facial images [6], [11].
Experimental behavior shows that ResNet50 is highly expressive but prone to over-
fitting when trained on smaller or less diverse datasets. This observation highlights
the importance of regularization and dataset diversity when deploying high-capacity
architectures for deepfake detection [5], [14].

3.2.4 MobileNetV2: Lightweight Architecture


MobileNetV2 achieves computational efficiency through depthwise separable convo-
lutions and inverted residual blocks. This design significantly reduces parameter
count and computational cost, making it suitable for resource-constrained environ-
ments [8], [11].
While MobileNetV2 offers strong efficiency, its lower representational capacity limits
its ability to detect subtle deepfake artifacts. Consequently, it provides a useful
lower-bound reference for architectural complexity in deepfake detection tasks.

16
Page 23 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 24 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

3.2.5 XceptionNet: Specialization in Spatial Artifacts


XceptionNet fully decouples spatial and channel-wise feature extraction, allowing
it to focus on localized spatial perturbations such as texture inconsistencies and
boundary artifacts [6], [11]. This specialization makes it particularly effective for
image-based deepfake detection.
However, due to its high expressive power, XceptionNet is more susceptible to over-
fitting, especially on smaller datasets or those with identity overlap. This reinforces
the necessity of explicit regularization strategies when using highly expressive archi-
tectures [5], [14].

3.2.6 Custom CNN Architecture Proposal


The proposed custom CNN architecture is designed to address the limitations ob-
served in pretrained models. Instead of prioritizing depth or parameter count, the
architecture emphasizes controlled capacity and regularization. Batch normalization
is employed to stabilize training, dropout layers reduce neuron co-adaptation, and
L2 regularization penalizes excessively large weights [11].
Additionally, a channel-wise attention mechanism is incorporated to prioritize facial
regions more likely to contain manipulation artifacts. This design aims to balance
expressiveness and generalization, enabling improved performance across datasets
with diverse characteristics [1], [15].

3.2.7 Dataset Description


Table 3.1 summarizes the datasets used in this phase along with their key charac-
teristics.
Table 3.1: Dataset Summary

Dataset Type Characteristics


140K Real and Fake Faces Image Large-scale, diverse identities
FaceForensics++ (C23) Image Compression artifacts
HardFake vs Real Faces Image High realism
RVF10K Image Balanced dataset

3.2.8 Model Comparison Strategy


All models are trained using identical preprocessing pipelines, fixed input resolu-
tions, and standardized evaluation metrics to ensure fairness and reproducibility.
This approach guarantees that observed performance differences arise from archi-
tectural design rather than experimental variation.

3.3 Trade-off Analysis and Architectural Justifi-


cation
An essential aspect of engineering-oriented research is the analysis of design trade-
offs. In deepfake detection, trade-offs exist between model complexity, detection ac-

17
Page 24 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 25 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

curacy, generalization ability, and computational feasibility [11], [14]. High-capacity


models such as ResNet50 and XceptionNet perform well on diverse datasets but are
more susceptible to overfitting. Lightweight architectures like MobileNetV2 gener-
alize better but lack sufficient expressiveness for subtle manipulations [5], [8].
EfficientNetB0 represents a compromise between stability and efficiency, while the
custom CNN architecture aims to balance capacity and regularization. These trade-
offs justify the inclusion of a custom-designed model alongside pretrained architec-
tures.

3.4 Overfitting Behavior and Dataset Sensitivity


Analysis
Overfitting is a critical challenge in deepfake detection due to repeated identities and
correlated artifacts in datasets [5], [14]. Deeper models tend to converge rapidly on
training data while exhibiting unstable validation performance on smaller datasets,
indicating reliance on dataset-specific cues.
In contrast, architectures with controlled capacity and explicit regularization demon-
strate more stable and predictable training behavior. The incorporation of dropout
and L2 regularization in the custom CNN reduces reliance on spurious correlations
and promotes learning of manipulation-relevant features.

3.5 Alternative Design Approaches


Alternative approaches, such as temporal modeling using 3D CNNs or recurrent
networks and transformer-based image detection, were considered but deferred to
later phases. These approaches introduce higher computational complexity and are
more suitable for Phase 3, which focuses explicitly on spatio-temporal modeling.

3.6 Simulation Environment and Experimental Con-


siderations
All experiments are conducted in cloud-based GPU environments using consistent
software frameworks, library versions, and training configurations. Fixed input reso-
lutions, standardized preprocessing, and uniform evaluation metrics are maintained
across models to minimize experimental variability and strengthen comparative con-
clusions.

3.7 Scope Limitations of the Present Design


Despite extensive evaluation, this phase has limitations. Frame-level image classifi-
cation does not capture temporal inconsistencies critical for advanced video-based
deepfake detection. Additionally, reliance on public datasets introduces potential
biases that may not fully represent real-world scenarios. These limitations motivate
the transition toward spatio-temporal modeling in subsequent phases.

18
Page 25 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 26 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

Chapter 4

Conclusion

This dissertation has in detail discussed the difficulties of identifying deepfake media
with deep-learning-based computer vision models, particularly the interactions be-
tween spatial (visual) and motion-based (temporal) visualization. However, errors
in the existing techniques, including facial manipulation vulnerability and some bias
to adversarial effects, continue to make these divisions ineffective in legal practice,
making it less efficient using these tools in practice . To fill these lapses, this thesis
suggests a new input-based detection claim through a transformer-based gatekeeper
framework to first sift through the bondage video frames and sequences by sup-
plying a dialogue framework to the principal system and apparatus of detection.
Through attention-based mechanisms, this gatekeeper would detect suspicious vi-
sual and motion patterns (e.g., texture anomalies, unnatural blinking) at personal
advantage above 95/97 percent at a robustly low latency (sub-20ms) on consumer-
tradeable GPUs. The following step as detailed in Chapter 1, entails labeling a
variety of real and manipulated videos, such as benign and adversarial videos, to
train and refine the gatekeeper model with the aid of underlying platforms such
as PyTorch, TensorFlow, and Hugging Face using the CUDA/cuDNN application.
Classification accuracy, F1-score, ROC-AUC, and inference time will be used after
a performance measure, and to facilitate the transparency of the applications, like
a digital forensics system, Grad-CAM visualization is to be used .

19
Page 26 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 27 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

Bibliography

[1] S. Das et al., Dynamic face augmentation and efficientnet/xception evaluation


for deepfake detection, arXiv Preprint, 2021.
[2] S. Das et al., Towards solving the deepfake problem: Improving detection using
dynamic face augmentation, arXiv:2102.09603, 2021.
[3] T. T. Nguyen et al., “Deep learning for deepfakes creation and detection,”
IEEE Transactions on Cybernetics, 2021.
[4] T. T. Nguyen et al., “Temporal modeling for deepfake detection using cnn and
lstm,” IEEE Transactions on Cybernetics, 2021.
[5] A. Malik et al., “Deepfake detection for human face images and videos: A
survey,” IEEE Access, 2022.
[6] A. Mitra et al., “Key video frame extraction for deepfake detection,” IEEE
Access, 2022.
[7] A. Mitra, S. P. Mohanty, P. Corcoran, and E. Kougianos, “A machine learning
based approach for deepfake detection in social media through key video frame
extraction,” IEEE Access, 2022.
[8] M. Nawaz et al., “Deepfake detection using error level analysis and deep learn-
ing,” ResearchGate, 2022.
[9] R. S. Ram et al., “Deep fake detection using computer vision-based deep neural
network with pairwise learning,” ResearchGate, 2022.
[10] B. Ghita et al., “Deepfake image detection using vision transformer models,”
Kaggle/University of Plymouth Dataset Study, 2023.
[11] A. Heidari et al., “Deepfake detection using deep learning methods: A sys-
tematic and comprehensive review,” Wiley Interdisciplinary Reviews (WIRES
Data Mining), 2023.
[12] A. Rahman et al., “Detection of deepfake using computer vision and deep
learning,” M.S. thesis, BRAC University, 2023.
[13] P. Z. Janakbari, “Detection and mitigation of deepfake attacks in cyberse-
curity: Leveraging computer vision and deep learning,” M.S. thesis, Theseus
Thesis Repository, 2024.
[14] A. Kaur et al., “Deepfake video detection: Challenges and opportunities,”
Springer AI Review, 2024.
[15] K. N. Ramadhani et al., “Improving video vision transformer for deepfake
video detection using facial landmarks and depthwise separable convolutions,”
IEEE Access, 2024.

20
Page 27 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642
Page 28 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

[16] H. Huang et al., “Sida: Social media image deepfake detection, localization
and explanation with large multimodal model,” in CVPR, 2025.

21
Page 28 of 28 - AI Writing Submission Submission ID trn:oid:::1:3473048642

You might also like