0% found this document useful (0 votes)
4 views7 pages

Deep Fake

This paper presents a method for detecting deepfake videos using a combination of ResNext Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) algorithms. The study shows that the model achieved a maximum accuracy of 90% with video data containing 40 and 60 frames, while lower accuracy was observed with 10 frames. The findings emphasize the importance of deepfake detection in maintaining media authenticity and preventing misinformation.

Uploaded by

janicebenita123
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views7 pages

Deep Fake

This paper presents a method for detecting deepfake videos using a combination of ResNext Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) algorithms. The study shows that the model achieved a maximum accuracy of 90% with video data containing 40 and 60 frames, while lower accuracy was observed with 10 frames. The findings emphasize the importance of deepfake detection in maintaining media authenticity and preventing misinformation.

Uploaded by

janicebenita123
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Deepfake Video Detection Using CNN-LSTM

Antonio Louie Jeffrey Harish Babu Selvam Madhan.K


UG Scholar UG Scholar Assistant Professor
Department of IT Department of IT Department of IT
St. Joseph’s College of Engineering, St. Joseph’s College of St. Joseph’s College of Engineering,
Engineering,
Chennai,India Chennai,India Chennai,India
antoniolouiejeffrey@[Link] harishselvam622@[Link] madhanckn@[Link]

Abstract – Deep-fake in videos is a video synthesis Neural Networks (CNN) and Long Short Term
technique by changing the people’s face in the video Memory(LSTM) classifiers.
with others’ face. Deep-fake technology in videos has
been used to manipulate information, therefore it is Keywords- Python, Convolutional Neural Networks
necessary to detect deep-fakes in videos. This paper (CNN) and Long Short-Term Memory (LSTM)
aimed to detect deep-fakes in videos using the classifiers
ResNext Convolutional Neural Networks (CNN) and
I. INTRODUCTION
Long Short-Term Memory (LSTM) algorithms. The
video data was divided into 4 types, namely video with Healthcare Deepfakes are images and videos, usually
10 frames, 20 frames, 40 frames and 60 frames. The created by deep neural networks to superimpose target
results of data classification showed that the highest subject’s face features over another in order to produce
accuracy value was 90% for data with 40 and 60 fake media. According to a recent report from the
frames. While data with 10 frames had the lowest Deeptrace Lab [1], there are almost 15,000 deefake
accuracy with 52% only. ResNext CNN-LSTM was media over the internet consisting of non obscene videos
able to detect deep-fakes in videos well even though targeting the politicians and the functioning of
the size of the image was [Link] fast evolvement of democratic societies. More than 13,000 deepfake videos
deepfake introduction technology is critically were found on different deepfake-specific porn sites [2],
threating media facts trustworthiness. The and about 96% of the deepfake content on web are
consequences impacting focused individuals and related topornography and mostly related to famous
establishments may be dire. In this work, we study celebrities to defame the individuals. Day by day,
the evolutions of deep studying architectures, deepfake’s findings are becoming so realistic that they
especially CNNs and Transformers. We diagnosed 8 are almost indistinguishable, and the substituted subjects
promising deep learning architectures, designed and are rigged to say things they never spoke [3]. Deepfake
evolved our deepfake detection fashions and methods have been extensively applied nowadays to
conducted experiments over well-installed deepfake produce enormous fake news, posing a serious threat to
datasets. Those datasets protected the modern 2nd communities worldwide, and have the potential to
and third generation deepfake datasets. We evaluated influence the masses as well as democratic and
the effectiveness of our evolved unmarried version geopolitical structure of a region [4]. Many organizations
detectors in deepfake detection and move datasets as well as private companies are investing heavily to
evaluations. This standardization is essential for the counter the challenges of deepfake. CNN with SVM
subsequent characteristic extraction technique, which Recurrent Neural Network(RNN) CNN with Long Short
makes a speciality of capturing important Term Memory (LSTM) [8], etc. Moreover, many
characteristics of the faces through local capabilities traditional ways have also been explored to detect
consisting of imply, widespread deviation, and manipulated media, such as exposing inconsistent in
variance. these statistical metrics provide a robust headposes, consideration of the background color
basis for information versions in facial attributes, manipulation [9]. However, most of these techniques are
enhancing the model’s potential to distinguish among focuses primarily on the spatial feature analysis and does
diverse identities that integrates Convolutional not include any temporal information. Since, most of the
deepfake media are in the form of video, therefore,
identifying inconsistencies in temporal information the necessity for continual learning in dynamic detection
(along with spatial inconsistencies) may enhance the environments. In a broader context, Gambín et al. [3]
classification accuracy of deepfakes. In this work, we presented a forward-looking survey exploring both
investigated a hybrid deep learning approach on current and emerging trends in deepfake technologies.
modelling the intra-frame as well as inter-frame features Their review spanned detection methods, attack vectors,
of videos to accurately identify its authenticity. and policy responses. A unique contribution of this work
Further,we incorporated a traditional temporal feature is the exploration of future threats, including the
analysis method, optical flow to help in extraction of the convergence of deepfakes with augmented reality and
temporal features. The optical flow implementation is voice synthesis. They advocate for a multipronged
based on to characterization of the motion of the defense strategy combining detection tools, public
subject’s face and the technique exploits the possible awareness, and platform-level safeguards. Almars [4]
inter-frame dissimilarities. A detail analysis has been focused on deep learning-based detection techniques,
carried out on to characterize the proposed model on offering a classification of methods such as autoencoders,
various sets of video data. The proposed model is CNNs, and RNNs. Their comparative analysis of model
evaluated based on its various performance parameters architectures and training data revealed a trade-off
such as Accuracy, Recall, Precision, F1-score, and AUC. between detection accuracy and computational
efficiency. This study served as a foundational resource
II. LITERATURE SURVEY for researchers entering the field, particularly in
understanding the evolution from handcrafted features to
The rapid evolution of deep learning and generative
end-to-end deep learning models. Gong and Li [5]
models has led to the proliferation of deepfakes,
extended the discussion by emphasizing the role of
synthetic media generated by algorithms such as GANs
datasets in benchmarking deepfake detection
(Generative Adversarial Networks). This phenomenon
performance. They cataloged various public datasets such
has raised significant concerns in media authenticity,
as FaceForensics++, Celeb-DF, and DFDC, highlighting
social trust, and security. Consequently, deepfake
their limitations in diversity, resolution, and real-world
detection has emerged as a critical research area within
complexity. Moreover, they reviewed deepfake detection
artificial intelligence, computer vision, and cybersecurity
algorithms across four categories: spatial-based,
domains. This literature study presents a comprehensive
frequency-based, temporal-based, and multimodal
review of recent contributions in deepfake detection from
techniques. Their work bridges the gap between theory
20 selected works, categorizing them into architectural
and practical implementation.
frameworks, ensemble models, surveys, dataset
benchmarking, and societal implications. Wazid et al. [1] Real-time detection has become increasingly relevant due
proposed a comprehensive framework focusing on the to the integration of deepfake detection systems into
architectural and security aspects of deepfake mitigation. consumer applications. Lanzino et al. [6] tackled this
The study emphasizes the importance of layered challenge by introducing a binary neural network (BNN)
architectures incorporating watermarking, blockchain, optimized for low-latency inference. Their model, named
and deep learning classifiers for robust authentication. ―Faster Than Lies,‖ leverages reduced precision
Furthermore, they discuss emerging challenges such as arithmetic to enable deployment on edge devices.
the lack of standardized datasets, ethical concerns, and Experimental evaluations showed that BNNs could detect
the increasing sophistication of generative models. Their manipulated videos with significant speedup and
work highlights the interdisciplinary nature of the marginal trade-offs in accuracy.
deepfake problem, urging collaboration between
legal,technical, and policy-making entities. Sharma et al. III . PROPOSED WORK
[2] addressed catastrophic forgetting in deepfake
detection by proposing a GAN-CNN ensemble model This research utilized two methods in the deepfake
integrated with generative replay mechanisms. Their detection process, the feature extraction process using
method preserves performance when exposed to evolving ResNext CNN. Furthermore, from the previous process,
fake media distributions. Experimental results on social information was obtained that can be used as a reference
media images demonstrated that their model could in the classification process with the LSTM method. The
maintain high accuracy over incremental training cycles, evaluation process was carried out on two ways, namely
outperforming traditional CNNs. This work underscores from the functional way of the application and the
performance of the methods used in the detection
process. Each of these testing methods were blackbox the confusion matrix method. Based on Figure 2 is
and confusion matrix. research stages.

A. Dataset

The data used in this study was digital image. In


deepfake detection, data wass obtained through data
provider sites that have been widely used by previous
researchers. The data would be used in making the
model, besides that it was also used to evaluate the build
system model. In general, research data was divided into
2, first was training data that was data to train and to Figure 2. Research stages
build models and second was data testing which was the
C. Planning the Model
data sets used in testing models after the training process
completed. The average length of all videos was 10
seconds with a standard frame rate of 30 frames per
second see in Figure 1.

Figure 3. Deepfake Model Planning

1) Preprocessing
Based on Figure 3 is Deepfake model planning.
Preprocessing is the initial stage in object recognition
by combining the concepts of digital image, pattern
recognition, mathematics, and statistics. In the
preprocessing stage, the captured image was processed
first to adjust in size and converted into a scale form
that was suitable for training data and trial data
focusing on the face. Frames without faces were
ignored. The final process in preprocessing stage was
storing the results of the faces that had been processed
in the previous stage. Each video had an average
duration of 10 seconds and was divided into 30 frames
Figure 1. Dataset in deepfake detection per second with a resolution of 112*112.

2) Deepfake Detection Model


B. Research stages
In this study, a combination of methods would be used
The research started by collecting data in the form of a when performing deepfake detection. The combined
video. The obtained data would go through the methods ware the ResNext CNN and LSTM methods.
preprocessing stage. This process changed the data into The feature extraction process would be carried out by
a dataset form that can be accepted by the system. And the CNN ResNext method while LSTM performed the
then all unnecessary things and noise in the video were classification. Data that has gone through the
removed. Only the required portion of the video were preprocessing process would be inputed into the model.
saved. The results of the preprocessing would be stored
and labeled as detected face markers. Furthermore, the
labeled data would be converted into a binary image a) ResNext Convolutional Neural Network (ResNext
with the feature extraction process using the ResNext CNN)
CNN method. After performing data extraction, the
data was classified into categories of fake data or real Convolutional Neural Network (CNN) is a
data. The final step would be to carry out the process of development of Multilayer Perceptron (MLP) designed
analyzing the accuracy of the detection process with to process two-dimensional data. CNN is included in
the type of Deep Neural Network because of its high
network depth and is widely applied to image data. In Accuracy
general, the convolutional layer detects edge features in
the image, then the subsampling layer would reduce
the dimensions of the features obtained from the Precision
convolutional layer, and finally passed them on to the
output node through the forward propagation process
[11][12][13].
Recall

(1) The Confusion Matrix contains all the raw information


ResNext CNN is one of the optimizations of CNN in about the predictions made by a classification model on
getting more optimal results. ResNext CNN is a a given data set. To evaluate the accuracy of model
method that has an advantage in maintaining generalizations, it is common to use test data sets that
computational complexity. It has good performance in are not used during the learning process of the model
reducing 6.7% errors in data processing [14]. This [20].
method has a simple design in extraction [15][16]. In
this study the ResNext CNN method was used in the
feature extraction process, the results of the extraction IV EXPERIMENTAL CONFIGURATION
would be classified using the LSTM method. Feature
Extraction process using ResNext CNN see in Figure In Deepfake detection in this study was carried out
4. using the ResNext CNN method with facial images
extraction taken from the frame. Frame captures were
done in stages with a 10-second video data set. Each
frame was intended to get frames with different facial
expressions or positions. In this study each video was
analyzed based on the number of frames. The number
of frames taken started from 10, 20, 40, 60. Each frame
had different results which would be categorized into,
true positive, false positive, false negative, true
Figure 4. Feature Extraction process using ResNext negative see in Figure 5, Figure 6 and Figure 7.
CNN

b) Long Short-Time Memory(LSTM)


LSTM algorithm is a component widely used in the
Recurrent Neural Networks (RNN). The algorithm can
rely on its learning sequence on pattern recognition
problems [17]. It has three states that help the network
to reduce long term dependency data. It has three steps,
forget step, input step and output step. The forget state
was used to remove redundant or useless data. The
input state functioned to process new data while the
output step was to process input data with cell status. Figure 5. True Positive Detector
The formulas of the three steps can be seen below [18]:

3) Model Evaluation
Classification algorithm performance can be measured
using a confusion matrix. There are several evaluation
indicators that can be used. In this research, the
indicators used were accuracy, precision, and recall.
Accuracy described the percentage of correct
classification of deep-fake video detection presented
follows[19]:
Figure 6. False Negative Detector
and recall values than data with a larger number of
frames, see in Table 3.

V. CONCLUSION

Figure 7. True Negative Detector Deep-fake video detection highly recommended to be


done to find out whether information is genuine or fake.
It can prevent mistakes in making decisions. Deep-fake
video classification was conducted by processing video
data, training the classification algorithm, evaluating, and
comparing the performance of the algorithm based on
video data on deep-fake video and real video. ResNext
CNN was used to process video data. In detecting deep-
fake videos, the videos should be extracted into an image
so that the image can be cropped to a smaller size, a size
of 100 x 100 pixels. This was to make the classification
process easier. LSTM was the model used to classify
image data and model performance needs to be tested
using indicators of accuracy, precision, and recall.
Classification results using ResNext CNN-LSTM
showed that the highest accuracy was 90%, precision
reached 100% and recall reached 97% for data from 60
frames. While data from 10 frames of video had lower
performance with 52% accuracy, 52% precision and 50%
recall. Based on the experimental results, it can be
Based on system testing on 100 tested data, the TP, TN, concluded that the combination of ResNext CNN and
FP, and FN values for each test were obtained, see in LSTM produced high performance even though the
Table 2. In a video with 10 data frames, the accuracy was image size was only 100 x 100 pixels. In addition,
52%, while precision and recall were 52% and 50% ResNext CNN-LSTM also had good performance on
respectively. This value was obtained from TP of 25, TN video data with 20 frames. For further research, it is
of 27, FP of 23 and FN of 25. better to explore the function of the LSTM to have more
optimal performance. Some researchers have also
Accuracy = explored certain facial features as datasets to include in
models, for example eyes, nose, ears, or mouth so it is
interesting to compare model performance between
Precision = models trained with the entire face and models trained
with partial facial features.
Recall =

Table 3. Table of Testing Confusion Matrix Classification Results (%)


Number of
Frames
Based on the classification results in Table 1, the Accuracy Precision Recall
performance of CNN and LSTM was lowest on video
data with 10 frames with 52% accuracy. The accuracy 10 Frames 52% 52% 50%
of the ResNext CNN and LSTM methods increased on
video with 20 frames reaching 88% and achieving 20 Frames 88% 88% 97%
stability on data video with 40 frames and 60 frames
with 90% accuracy. It was similar with the value of 40 Frames 90% 90% 97%
precision and recall. Deepfake classification for data
with 10 frames indicated lower accuracy, precision, 60 Frames 90% 100% 97%`
REFERENCES [11] I. A. Anjani, Y. R. Pratiwi, and S. Norfa Bagas
Nurhuda, “Implementation of Deep Learning Using
[1] A. Brunetti, D. Buongiorno, G. F. Trotta, and V.
Bevilacqua, “Computer vision and deep learning Convolutional Neural Network Algorithm for
Classification Rose Flower,” J. Phys. Conf. Ser., vol.
techniques for pedestrian detection and tracking: A 1842, no. 1, 2021.
survey,” Neurocomputing, vol. 300, pp. 17–33, 2018.
[12] Q. Zhang, M. Zhang, T. Chen, Z. Sun, Y. Ma, and
[2] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. B. Yu, “Recent advances in convolutional neural network
Wojna, “Rethinking the Inception Architecture for acceleration,” Neurocomputing, vol. 323, pp. 37–51,
Computer Vision,” Proc. IEEE Comput. Soc. Conf. 2019.E-ISSN 2548-7779 ILKOM Jurnal Ilmiah Vol. 14, No. 3,
Comput. Vis. Pattern Recognit., vol. 2016-Decem, pp. December 2022, pp.178-185 185 Abidin, et. al. (Deepfake
2818–2826, 2016. detection in videos using Long Short-Term Memory and CNN
ResNext)
[3] H. Wang, Y. Wang, Z. Zhou, X. Ji, D. Gong, and J.
Zhou, “CosFace: Large Margin Cosine Loss for Deep [13] M. A. Saleem, N. Senan, F. Wahid, M. Aamir, A.
Face Recognition,” Cvpr, pp. 5265–5274, 2018. Samad, and M. Khan, “Comparative Analysis of Recent
Architecture of Convolutional Neural Network,” Math.
[4] O. Victoria and I. P. Solihin, “Pendeteksi wajah Probl. Eng., vol. 2022, 2022.
secara realtime menggunakan Metode Eigenface,”
[14] K. Annapurani and D. Ravilla, “CNN based image
SEINASI-KESI (Seminar Nas. Inform. Sist. Inf. Dan classification model,” Int. J. Innov. Technol. Explor.
Keamanan Siber), pp. 126–131, 2018. Eng., vol. 8, no. 11 Special Issue, pp. 1106–1114, 2019.

[5] D. Pan, L. Sun, R. Wang, X. Zhang, and R. O. [15] S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He,
Sinnott, “Deepfake Detection through Deep Learning,” “Aggregated residual transformations for deep neural
Proc. - 2020 IEEE/ACM Int. Conf. Big Data Comput.
Appl. Technol. BDCAT 2020, pp. 134–143, 2020. networks,” Proc. - 30th IEEE Conf. Comput. Vis. Pattern
Recognition, CVPR 2017, vol. 2017-Janua, pp. 5987–
[6] M. Westerlund, “The emergence of deepfake 5995, 2017.
technology: A review,” Technol. Innov. Manag. Rev.,
vol. 9, no. 11, pp. 39–52, 2019. [16] G. Pant, D. P. Yadav, and A. Gaur, “ResNext
convolution neural network topology-based deep learning
[7] S. T. Suganthi et al., “Deep learning model for deep
fake face recognition and detection,” PeerJ Comput. model for identification and classification of
Pediastrum,” Algal Res., vol. 48, no. April, p. 101932,
Sci., vol. 8, pp. 1–20, 2022. 2020.

[8] S. Agarwal, H. Farid, T. El-Gaaly, and S. N. Lim, [17] P. N. Srinivasu, J. G. Sivasai, M. F. Ijaz, A. K. Bhoi,
“Detecting Deep-Fake Videos from Appearance and W. Kim, and J. J. Kang, “Classification of skin disease
using deep learning neural networks with mobilenet v2
Behavior,” 2020 IEEE Int. Work. Inf. Forensics Secur. and lstm,” Sensors, vol. 21, no. 8, pp. 1–27, 2021.
WIFS 2020, 2020.
[18] C. I. Garcia, F. Grasso, A. Luchetta, M. C. Piccirilli,
[9] S. Megawan, W. S. Lestari, and A. Halim, “Deteksi L. Paolucci, and G. Talluri, “A comparison of power
Non-Spoofing Wajah pada Video secara Real Time quality disturbance detection and classification methods
using CNN, LSTM and CNN-LSTM,” Appl. Sci., vol. 10,
Menggunakan Faster R-CNN,” vol. 3, no. 3, 2022.
no. 19, pp. 1–22, 2020.
[10] P. Ranjan, S. Patil, and F. Kazi, “Improved
[19] X. Li, S. Li, J. Li, J. Yao, and X. Xiao, “Detection of
generalizability of deep-fakes detection using transfer
fake-video uploaders on social media using Naive
learning
Bayesian model with social cues,” Sci. Rep., vol. 11, no.
based CNN framework,” Proc. - 3rd Int. Conf. Inf. 1, pp. 1–11, 2021.
Comput. Technol. ICICT 2020, pp. 86–90, 2020.
[20] O. Caelen, “A Bayesian interpretation of the
confusion matrix,” Ann. Math. Artif. Intell., vol. 81, no.
3–4, pp. 429–450, 2017.

You might also like