Image Forgery Detection with ViTs
Image Forgery Detection with ViTs
[Link]
REGULAR PAPER
Abstract
Due to the easy availability of software over the Internet, any naive user can tamper the images for entertainment purposes
or to defame a personality by circulating over social media networks. The practice of image tampering is a serious issue and
can attract legal action if proven guilty. Forensic researchers employ various methods to detect and localize image forgeries.
In this research, we use a Vision transformer (ViT) as a method for binary classification of images distinguishing forged
and unforged images. Further, we use a pre-trained Segment Anything Model(SAM) which is fine-tuned with custom data to
adaptively recognize patterns indicating forged regions within the images. SAM can localize these forged areas and is leveraged
to create templates by extracting the identified regions. The proposed method is rigorously tested across various datasets,
including CASIA v1.0, CASIA v2.0, MICC-F2000, MICC-F600, and Columbia. Through comprehensive experimentation,
our approach showcases considerable promise yielding accuracy in image forgery classification and localization. Our model’s
robustness and adaptability make it an attractive tool for forensic analysis in diverse scenarios, contributing to the advancement
of multimedia forensics security research.
Keywords Image forensics · Forgery detection · Forgery localization · Vision transformer · Segment anything model ·
Binary mask
123
International Journal of Multimedia Information Retrieval (2025) 14:8 Page 3 of 11 8
3 Related work
Segment Anything Model (SAM), is a powerful deep-
learning architecture designed specifically for image seg- Bi et al. [4] provides a thorough overview of Evolutionary
mentation tasks. It leverages a novel transformer-based Computation (EC) methods used in computer vision and
encoder-decoder structure to achieve state-of-the-art results image processing. The survey delves into the noteworthy
in various segmentation applications. The Segment Any- accomplishments of EC in tackling image-related problems
thing Model (SAM), introduced by Kirillov et al. [15], is including edge detection, picture segmentation, feature anal-
a novel model for image segmentation. In addition, the ysis, classification, and object detection. The authors provide
authors have created the largest segmentation dataset to date, insights into the history, current state, and future directions
with over 1 billion masks extracted from 11 million autho- of Evolutionary Computer Vision (ECV) by discussing the
rized and privacy-respecting photos by utilizing an effective contributions of several EC techniques. This survey is a great
model within a data-collecting loop. The Segment Anything resource for learning about the state of EC in the context of
Model (SAM) can move zero-shot to new image distribu- computer vision and image analysis.
tions and tasks. The authors assessed SAM’s performance Han et al. [10] provide a two-stage CNN transfer learning
on a range of tasks and found that it frequently matches or and web data augmentation approach for image categoriza-
outperforms previous fully supervised approaches in terms tion. This method effectively transfers pre-trained network
of zero-shot performance. Integrating the Segment Anything features, which help deep CNNs overcome the problem of
Model (SAM) with Focal Loss for accurate identification having insufficient training data. The strategy efficiently
and localization of forged regions in images requires Model addresses overfitting on small datasets and minimizes the
Selection & Adaptation, Improving Localization Accuracy, requirement for large amounts of training data by supple-
Computational Considerations, and Evaluation metrics. The menting the original dataset with images from the Internet.
SAM dataset (SA-1B), which includes 1 billion masks and Fine-tuning and hyper-parameter tuning is done via Bayesian
11 million images is publicly available to support future optimization. Experiments conducted on six small datasets
research in foundation models for computer vision. Integrat- showed better performance, especially with ResNet surpass-
ing the Segment Anything Model (SAM) with Focal Loss for ing the performance of the most advanced models.
accurate identification and localization of forged regions in Mehrjardi et al. [18] examine deep learning-based image
images requires Model selection & adaptation, Enhancing fraud detection in their survey, highlighting the growing sig-
SAM with Focal Loss, Training pipeline & dataset aug- nificance of image authenticity and reliability especially in
mentation, Improving localization accuracy, Computational applications related to medicine, forensics, and the legal sys-
considerations, and Evaluation & metric. Figure 4 shows the tem. The survey examines several forms of image forgeries,
architecture of the Segment Anything Model (SAM). evaluation criteria, benchmark datasets, and conventional
123
8 Page 4 of 11 International Journal of Multimedia Information Retrieval (2025) 14:8
techniques for detecting forgeries. While conceding the forgery detection were found by the enhanced SIFT descrip-
shortcomings of older methodologies, the authors underline tor.
the robustness and automatic recognition capabilities of deep In their extensive review, Han et al. [11] provide an in-
learning methods, placing them as a popular alternative in depth synopsis of the developments in ViT architecture and
computer vision. address the important topics including training methods,
The review paper by Taneja et al. [24] examines digital applications, and architecture design, providing insightful
picture anti-forensics and highlights how it affects forensic information about how the Vision Transformers market is
difficulties. The article discusses artifacts like JPEG block- developing. This survey is a great way to learn about the
ing and contrast enhancement that are essential for forensic current state of the art in virtual reality research.
methods. By removing traces of their activities, anti-forensic Asiri et al. [2] provides a thorough investigation into
techniques make detection more difficult. The review also the classification of brain tumors utilizing fine-tuned vision
describes attacks such as JPEG anti-forensics that help foren- transformers (ViTs). Several pre-trained ViT models (R50-
sic analysts create reliable forgery detection for applications ViT-l16, ViT-l16, ViT-l32, ViT-b16, and ViT-b32) are used in
like cybercrimes by summarizing contemporary methodolo- the study to improve the state-of-the-art in brain tumor clas-
gies. sification. With 4855 training images and 857 testing images,
He et al. [12], introduced a unique method titled "Image ViT-b32 was found to be the best-performing model trained
Copy-Move Forgery Detection via Deep Cross-Scale Patch- on four different tumor classes. The outcomes outperform
Match", which addresses the practical limitations of existing current techniques, demonstrating the usefulness of the sug-
deep learning algorithms. By combining traditional and gested strategy and setting a standard for further study on the
deep learning techniques [20, 21], end-to-end deep cross- categorization of brain tumors.
scale patch-match is introduced and designed specifically The difficulties of accurately detecting and localizing
for copy-move forgery detection (CMFD). The method face forgeries due to lack of pixel-by-pixel supervision are
stresses explicit point-to-point matching using features from tackled by well-trained Segment-Anything Model(SAM) in
high-resolution scales, boosting generalizability to different Detect Any Deepakes(DAD) framework [16]. DAD frame-
copy-move contents. Comprehensive trials show that the sug- work effectively captures both short and long-range forging
gested framework performs better than the current methods. contexts with the help of the Multiscale Adapter and the
A novel CrossViT model was developed by Chen et al. Reconstruction Guided Attention (RGA) module, increasing
[6] which is a type of vision transformer (ViT) that investi- the sensitivity of the model to forged regions.
gates the multi-scale feature representations in transformer In the work titled "Segment Anything", Kirillov et al. [15]
models for image classification. The dual-branch transformer introduced the Segment Anything (SA) project. For image
uses attention mechanisms to fuse images by processing segmentation, the project presents a new task, model, and
image patches of different widths. A cross-attention-based dataset. To create the largest segmentation dataset to date-
token fusion module is proposed to improve model efficiency, which includes over 1 billion masks on 11 million licensed,
which results in a considerable reduction in computation time privacy-respecting images-the authors employed an effective
when compared to quadratic computation. Numerous exper- model in the data gathering loop. Segment Anything Model
imental tests demonstrate that CrossViT performs better than (SAM) is a model that can be zero-shot transferred to new
current works and CNN models. image distributions and jobs since it is promptable by design
In the thorough analysis of attention mechanisms in com- and training. On a variety of tasks, the zero-shot performance
puter vision, Guo et al. [9] examined how well the attention of SAM is shown to be impressive and frequently on par with
mechanisms work for various tasks, including semantic seg- or better than previous fully supervised results.
mentation, object recognition, and image classification. The
attention mechanisms dynamically modify weights based on
input image elements, drawing inspiration from the human
visual system. The survey provides a comprehensive analysis
of attention divided into methods like temporal, geograph- 4 Challenges and contributions
ical, branch, and channel and also suggests future paths of
inquiry. Overall, the major challenges encountered in optimizing
A feature-based copy-move forgery detection (CMFD) models for classification and localization were instrumental
technique was introduced by Bin Yang et al. [28] in which in refining the proposed methodologies. The contributions
key points are found using a modified Scale Invariant Fea- lie in the successful implementation of ViT-l32 for supe-
ture Transform (SIFT) detector. To distribute the key points rior image classification and the introduction of SAM with
throughout the image, a key-point distribution approach focal loss for accurate object localization, addressing critical
was developed. In the end, the critical spots for copy-move aspects in the field of forensic image analysis.
123
International Journal of Multimedia Information Retrieval (2025) 14:8 Page 5 of 11 8
123
8 Page 6 of 11 International Journal of Multimedia Information Retrieval (2025) 14:8
where pt is the predicted probability, and γ is the focus- Table 2 Details of datasets used
ing parameter [17]. Dataset Authentic Forged Type Mask
2. Dice Focal Loss: The Dice Focal Loss is defined as:
CASIA v1.0 800 459+462 CM,SP Yes
CASIA v2.0 7491 5123 CM,SP Yes
Dice Focal Loss = (1 − Dice Coefficient)γ · Focal Loss
MICC-F2000 1300 700 CM No
MICC-F600 440 160 CM Yes
where the Dice Coefficient is used in conjunction with
Columbia 183 180 SP No
the Focal Loss [19].
3. Adam Optimizer: The Adam Optimizer is an adaptive
learning rate optimization algorithm. The update rule is
given by: by a wide array of copy-move image forgery instances, ulti-
mately advancing the field of image forensics.
m̂ t MICC-F600 introduces complexity in the realm of copy-
θt+1 = θt − lr · move image forgery detection. With a collection of 600
v̂t + challenging images, this dataset has intricate geometric trans-
formations and introduces additional noise. Researchers turn
where θt is the parameter at time t, m̂ t and v̂t are the bias- to MICC-F600 to test the robustness and efficacy of their
corrected moving averages of the gradient and its square, algorithms under more demanding conditions, pushing the
lr is the learning rate (lr = 1 × 10−5 in the provided boundaries of copy-move image forgery detection capabili-
program), and is a small constant to prevent division ties.
by zero [14]. Columbia The Columbia dataset emerges as an extensive
resource, comprising copy-move forgery images, splicing
images, and a set of pristine images. With a total of 3876
6 Results and analysis images, this dataset provides a diverse array of forgery and
non-forgery scenarios. Researchers leverage the Columbia
6.1 Dataset dataset for training and testing purposes, benefitting from
its inclusion of various image manipulation techniques. This
CASIA 1.0 stands as a pivotal dataset in the realm of diversity contributes to the dataset’s significance in advanc-
image forgery detection, focusing on copy-move forgery ing the field of image forensics. Table 2 shows details of the
images. Comprising a substantial collection of 4000 images, datasets used.
this dataset offers a diverse array of realistic scenarios, Figure 5 illustrates samples from the datasets, showcasing
showcasing varying degrees of manipulation. The images two images identified as forged and one as authentic. This
within CASIA 1.0 are sourced from different origins and visual representation provides a glimpse into the diversity of
exhibit transformations such as rotation, scaling, and bright- the dataset and the varying characteristics between forged
ness adjustments. Researchers widely employ this dataset to and authentic images shown in Table 2.
assess and enhance the effectiveness of algorithms in detect-
ing copy-move forgery within digital images. 6.2 Computational metrics
CASIA 2.0 extends the complexity of image forgery
scenarios. This dataset exclusively consists of copy-move 6.2.1 Precision
forgery images, totaling 5000 in number. Notably, CASIA
2.0 introduces more intricate operations, including multi- Precision measures the accuracy of positive predictions made
ple copy-moved regions, geometric transformations, and the by the model. It is calculated as the ratio of true positive
intentional addition of noise. The increased complexity chal- predictions to the sum of true positives and false positives.
lenges of image forgery detection algorithms, make CASIA A high precision indicates that the model has fewer false
2.0 a valuable resource for researchers striving to improve positives.
the robustness and accuracy of their methodologies.
MICC-F2000 Claiming the title of the largest publicly True Positives
Precision =
available dataset for copy-move forgery detection, MICC- True Positives + False Positives
F2000 is a rich resource comprising 2000 images. This
dataset spans a diverse spectrum of complexity levels and 6.2.2 Recall (sensitivity or true positive rate)
manipulation scenarios, making it a comprehensive bench-
mark for algorithm evaluation. Researchers engage with Recall measures the ability of the model to capture all positive
MICC-F2000 to explore and address the challenges posed instances. It is calculated as the ratio of true positive predic-
123
International Journal of Multimedia Information Retrieval (2025) 14:8 Page 7 of 11 8
Intersection Area
IoU =
Union Area
6.3 Results
123
8 Page 8 of 11 International Journal of Multimedia Information Retrieval (2025) 14:8
Fig. 6 Forged image from CASIAv1.0 Dataset, with Predicted Mask Fig. 7 Forged image from CASIAv2.0 Dataset, with Predicted Mask
and Probability Map and Probability Map
123
International Journal of Multimedia Information Retrieval (2025) 14:8 Page 9 of 11 8
For instance, on CASIA v1.0, after 20 epochs, the mean 6.4 Result comparison
IOU reached 0.9052. Subsequently, employing Focal Loss
yielded promising results, surpassing Dice Focal Loss in In our experiments, the ViT-l32 model outperformed other
certain scenarios. For instance, on CASIA v1.0, after 10 Vision Transformer configurations like R50-ViT-l16, ViT-
epochs, the mean IOU was 0.9089. These results collec- b16, ViT-l16, and ViT-b32 due to its superior feature repre-
tively underscore the efficacy of our chosen model and sentation and deeper architecture. With a larger patch size
methodology in both classification and localization tasks, and enhanced capability to capture global context, it can
contributing valuable insights to the broader field of deep detect subtle image forgeries more comprehensively than
learning for image analysis. Ensuring an algorithm performs ever before, achieving training accuracy of 97.5% and vali-
reliably on complex datasets with intricate transformations dation accuracy of 94.5% on the CASIA v2.0 dataset, which
123
8 Page 10 of 11 International Journal of Multimedia Information Retrieval (2025) 14:8
Xu et al. [27] CASIA v1.0 0.816 Funding No funding is available for the Article Processing Charge (or
Tan et al. [23] CASIA v2.0 0.0.8394 publication fee).
Ganapathi et al. [8] CASIA v2.0 0.8816 Data Availability The authors confirm that the data supporting the
– MICC-F2000 – findings of this study are available within the article [and/or] its supple-
– MICC-F600 – mentary materials.
Ganapathi et al. [8] Columbia 0.9056 Code availability Available with the authors and provided to the
Proposed CASIA v1.0 0.9089 research community based on request and mutual understanding.
CASIA v2.0 0.9371
MICC-F2000 0.8938 Declarations
MICC-F600 0.8507
Columbia 0.9179 Conflict of interest There are no Conflict of interest.
References
served as the basis for future experiments. Table 5 shows 1. Amiri E, Mosallanejad A, Sheikhahmadi A et al Copy-move
the results of image classification on the proposed model forgery detection using eoa, dwt and dct. Pamukkale Univ J Eng
compared rigorously against state-of-the-art methods. Mov- Sci, 1000(1000):0–0
ing towards object localization, we employed the Segment 2. Asiri AA, Shaf A, Ali T, Pasha MA, Aamir M, Irfan M, Alqahtani S,
Alghamdi AJ, Alghamdi AH, Alshamrani AFA (2023) Advancing
Anything Model (SAM) with Focal Loss for varying epochs. brain tumor classification through fine-tuned vision transformers:
In Table 6, the proposed model is shown to be superior to a comparative study of pre-trained models. Sensors 23(18):7913
existing methods in terms of IOU. 3. Bayram S, Sencar HT, Memon N (2009) An efficient and robust
In summary, our proposed model consistently outper- method for detecting copy-move forgery. In: 2009 IEEE Interna-
tional Conference on Acoustics, Speech and Signal Processing, pp.
formed existing methods in image classification and show-
1053–1056. IEEE
cased superior object localization capabilities across diverse 4. Bi Y, Xue B, Mesejo P, Cagnoni S, Zhang M (2022) A survey on
datasets, affirming its efficacy in forensic image analysis. evolutionary computation for computer vision and image analysis:
123
International Journal of Multimedia Information Retrieval (2025) 14:8 Page 11 of 11 8
past, present, and future trends. IEEE Trans Evolutionary Comput 19. Milletari F, Navab N, Ahmadi SA (2016) V-net: Fully convolutional
27(1):5–25 neural networks for volumetric medical image segmentation. In:
5. Chen B, Qi X, Zhou Y, Yang G, Zheng Y, Xiao B (2020) Image 2016 fourth international conference on 3D vision (3DV), pp. 565–
splicing localization using residual image and residual-based fully 571. IEEE
convolutional network. J Vis Commun Image Represent 73:102967 20. Patil G, Palaiahnakote S, Gornale SS, Lopresti DP (2024) Altered
6. Chen CF (Richard), Fan Q, Panda R (2021) Crossvit: Cross- handwritten text detection in document images using deep learning.
attention multi-scale vision transformer for image classification. In: Int J Pattern Recognit Artific Intell 38(03):2452006
Proceedings of the IEEE/CVF International Conference on Com- 21. Patil G, Shivakumara P, Gornale SS, Pal U, Blumenstein M (2023)
puter Vision (ICCV), pp. 357–366 A new robust approach for altered handwritten text detection. Mul-
7. Deng L, Peng J, Deng W, Liu K, Cao Z, Wang W (2022) A dual- timed Tools Appl 82(14):20925–20949
stream input faster-cnn model for image forgery detection. In: 22. Shi YQ, Chen C, Chen W (2007) A natural image model approach
International Conference on Mobile Networks and Management, to splicing detection. In: Proceedings of the 9th workshop on Mul-
pp. 105–115. Springer timedia & security, pp 51–62
8. Ganapathi II, Javed S, Ali SS, Mahmood A, Vu N-S, Werghi N 23. Tan Y, Li Y, Zeng L, Ye J, Li X, et al (2023) Multi-scale target-aware
(2022) Learning to localize image forgery using end-to-end atten- framework for constrained image splicing detection and localiza-
tion network. Neurocomputing 512:25–39 tion. arXiv preprint arXiv:2308.09357
9. Guo M-H, Xu T-X, Liu J-J, Liu Z-N, Jiang P-T, Mu T-J, Zhang S-H, 24. Taneja N, Bramhe VS, Bhardwaj D, Taneja A (2023) Understand-
Martin RR, Cheng M-M, Hu S-M (2022) Attention mechanisms in ing digital image anti-forensics: an analytical review. Multimedia
computer vision: a survey. Comput Vis Media 8(3):331–368 Tools and Applications, pp. 1–22
10. Han D, Liu Q, Fan W (2018) A new image classification method 25. Tyagi S, Yadav D (2023) Forensicnet: modern convolutional neural
using cnn transfer learning and web data augmentation. Expert Syst network-based image forgery detection network. J Forensic Sci
Appl 95:43–56 68(2):461–469
11. Han K, Wang Y, Chen H, Chen X, Guo J, Liu Z, Tang Y, Xiao A, 26. Xiao B, Wei Y, Bi X, Li W, Ma J (2020) Image splicing forgery
Xu C, Xu Y, Yang Z, Zhang Y, Tao D (2023) A survey on vision detection combining coarse to refined convolutional neural network
transformer. IEEE Trans Pattern Anal Mach Intell 45(1):87–110 and adaptive clustering. Inf Sci 511:172–191
12. He Y, Li Y, Chen C, Li X (2023) Image copy-move forgery detec- 27. Xu Y, Muhammad I, Aiqing F, Jiangbin Z (2023) Multi-scale
tion via deep cross-scale patchmatch. In: 2023 IEEE International attention network for detection and localization of image splicing
Conference on Multimedia and Expo (ICME), pages 2327–2332. forgery. IEEE Trans Instrum Measur
IEEE 28. Yang B, Sun X, Guo H, Xia Z, Chen X (2018) A copy-move
13. Kadam KD, Ahirrao S, Kotecha K et al (2022) Efficient approach forgery detection method based on cmfd-sift. Multimed Tools Appl
towards detection and identification of copy move and image splic- 77:837–855
ing forgeries using mask r-cnn with mobilenet v1. Comput Intell 29. Zhang J, Wang H, He P (2023) Dual-branch multi-scale densely
Neurosci connected network for image splicing detection and localization.
14. Kingma DP, Ba J (2014) Adam: A method for stochastic optimiza- Signal Process Image Commun 119:117045
tion. arXiv preprint arXiv:1412.6980
15. Kirillov A, Mintun E, Ravi N, Mao H, Rolland C, Gustafson L, Xiao
T, Whitehead S, Berg Alexander C, Lo WY et al (2023) Segment
Publisher’s Note Springer Nature remains neutral with regard to juris-
anything. arXiv preprint arXiv:2304.02643
dictional claims in published maps and institutional affiliations.
16. Lai Y, Luo Z, Yu Z (2023) Detect any deepfakes: Segment any-
thing meets face forgery detection and localization. In: Chinese
Springer Nature or its licensor (e.g. a society or other partner) holds
Conference on Biometric Recognition, pp. 180–190. Springer
exclusive rights to this article under a publishing agreement with the
17. Lin TY, Goyal P, Girshick R, He K, Dollár P (2017) Focal loss for
author(s) or other rightsholder(s); author self-archiving of the accepted
dense object detection. In Proceedings of the IEEE international
manuscript version of this article is solely governed by the terms of such
conference on computer vision, pp. 2980–2988
publishing agreement and applicable law.
18. Mehrjardi FZ, Latif AM, Zarchi MS, Sheikhpour R (2023) A survey
on deep learning-based image forgery detection. Pattern Recogni-
tion, pp. 109778
123