Project Report 1
Project Report 1
Project Report
on
Detection of AI Plagiarism in Research
Papers in Real world submitted as
partial fulfillment for the award of
BACHELOROFTECHNOLOGY DEGREE
SESSION 2025-26
DECLARATION
We hereby declare that this submission is our own work and that, to the best of our
knowledge and belief, it contains no material previously published or written by
another person nor material which to a substantial extent has been accepted for the
award of any other degree or diploma of the university or other institute of higher
learning, except where due acknowledgment has been made in the text.
Signature
Name(s):
Roll
No.(s):
Date:
ii
CERTIFICATE
Date:
ACKNOWLEDGEMENT
It gives us a great sense of pleasure to present the report of the B. Tech Project
undertaken during B. Tech. Final Year. We owe special debt of gratitude to Dr.
Date:
Sig. (s):
Name
(s):
Roll No.:
ABSTRACT
Academic integrity sits at the heart of real education and honest research.
These days, with information just a click away, plagiarism isn't as simple as
copying text. Now it's about clever paraphrasing and twisting ideas Most
oldschool plagiarism checkers stick to matching words, phrases, or chunks of
text. Sure, they catch straight-up copying, but they struggle when someone
rewrites or rearranges the original. Plus, a lot of the commercial tools out
there cost too much, keep their tech locked away, and don't even give students
useful feedback. The result? We need something smarter, easier to use, and
actually helpful for learning. That's why this research introduces an
Alpowered plagiarism detection system that looks beyond the words on the
page. Instead of just matching vocabulary, it uses Google Gemini 2.5 Flash
and live web grounding to get what the text actually means. It checks context,
intent, and how ideas fit together. The system picks up on plagiarism even if
someone has completely reworded or reshaped the content. Everything runs in
your browser-thanks to [Link] and [Link]-so there's nothing heavy to install.
Your data stays private since everything happens locally and through secure
APIs. You can type in text directly or upload a PDF, whichever works for you.
But here's the real game-changer: the platform doesn't just slap a "plagiarized"
label on your work.
Contents
Introduction 1
Detecting Deep Fake Text . . . . . . . . . . . . . . . . . . . . . . . . . 1 Challenges with Deep
Detectors . . . . . . . . . . . . . . . . . . . . . . 1 Goals. . . . . . . . . . . . . . . . . . . . . . . . . . .
. . . . . . . . . 2 Contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2
Thesis Outline ............................... 3
Background 4
History & context . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 The State of the
Literature . . . . . . . . . . . . . . . . . . . . . . . . 5 About BERT Transformers . . . . . . . . .
............... 5
RoBERTa Usage in this text . . . . . . . . . . . . . . . . . . . . . . . . 6
Bibliography 30
A Attachments 39
A.1 Digital attachments . . . . . . . . . . . . . . . . . . . . . . . . . . 39
A.1.1 GitHub repos . . . . . . . . . . . . . . . . . . . . . . . . . 39
A.1.2 Datasets used . . . . . . . . . . . . . . . . . . . . . . . . . 39
Vita 40
Introduction
Detecting Deep Fake Text
Originally, the plan was to make models detecting video deep fakes by sequencing
CNN frames of video and making some novel contribution. However, the academic
body on vision was quite saturated and is/was difficult to contribute with true novelty.
In preparation for the literature review of the thesis, natural language processing was
chosen instead.
In replacement of detecting deep fake video, the idea was to detect deep fake text.
With the switch, the number of research papers was comparatively quite small. For
the SLR, there are only 50 papers in review because of this size. This SLR too, is the
first of its kind in regards to NLP. Also, the papers inspired by this topic are unique and
make a good contribution.
The motivation, as it had been, was to train a neural network to detect fake text.
The proposed papers were in regards to creating detectors and attacking them.
Though after the publication of Chat-GPT many more papers were added to the body
quite quickly. Then, to make a more meaningful contribution the topics were changed
to mutation, with the SLR coming in just in time as novel research before another
literature review on detectors could be made.
1
future steps mentioned in the below literature review are now being enacted. Though
the SLR is still relevant, the future steps are being filled in. It easy to write a paper
which at one point is highly important and relevant to then be outdated a year later.
Still, there are so much open research avenues, it will probably take some time to be
saturated.
Goals
The goal of this thesis is to make future work easier to accomplish for this domain.
This is done by reviewing the state of the literature in regards to detecting fake text
and to contribute to the literature by proposing a solution to mutation based
adversarial attacks. Because in mid 2022 the body of research was so small, the review
was the first of its kind in this domain. The review of the literature was made using
the PRISMA methodology to sift down from a little over 1000 NLP research papers in
search of those machine-centric text detector journals.
As well, chapter 3 on mutations, is an extension of Wolff [2020] where we used
the COCO Image captioning dataset to create a potential solution to the transformer
mutation problem. With these motivations set, the first order of business was to
choose architectures for the given experiments. With the advice of Dr. Liang and Dr.
Alsmadi, the proposed architecture was transformers. For the entirety of research
RoBERTa was used to label sequenced data/text in a classical transformer fashion.
The idea is to mutate text to appear as human with small changes to words or
letters. The most damaging component of mutation appeared to be using a
vocabulary not recognized by the transformer. Automatically, the transformer took
foreign vocabulary as human text and not synthetic. The research was based on this
issue, which when published was quite novel.
Contributions
In this thesis, a snapshot of the literature was made and a workflow for beating
mutation adversarial attacks was proposed with some example mutation operators.
These contributions allow future authors to find new paths of research both for
detecting fake text and attacking them with mutations. The main contributions of this
thesis are:
• An overarching view of the research done in 2022: The finding from the
literature review for this research domain is its status of being under
researched.
2
• Trends and statistics of the literature: The statistic being increasing growth,
predicted to increase much more in mid 2022. This help up to be true with the
publicity to transformers in public media (GPT3.5)
• A list of weaknesses the current overall research has: These will be the same
as research avenues. There are varying degrees as to which things are under
researched, so that is written out in the SLR.
Thesis Outline
The rest of this document consists of the following:
Background
History & Context
Since the creation of the perceptron in the late 1950’s, neural networks have been
used as a theoretical model for machine learning but had been limited by the
computational proficiency of our machines. Over the past three decades, increasing
computational power has allowed neural network research to flourish at an
unprecedented rate, not due to mathematical limitations, but by the processing
ability of our machines. It all started with the perceptron, a larger layered
mathematical formula, which processes numbers forward:
3
Figure 1.1: A basic perceptron
This above is a neural network and is typically much larger and grand. Numbers
moving from left to right are known as “feeding forward” and from right to left, “back
propagation”. These concepts are old and dated, with math dating back to the 60s,
70s & 80s. Back propagation is the most complex of the two, requiring calculus to
reverse the process in a “predictive” way.
In the 90s and early 2000s more research had been completed on these networks
and with the advent of smaller and more powerful processing chips researchers were
able to implement these mathematical models more commonly and with that more
existing types of models were finally made as software.
Examples of neural networks include CNN (Convolutional Neural Networks), RNN
(Recurrent Neural Networks) and Transformers in that chronological order. These, of
course, have their own use cases due to how, mathematically, their output layer is
calculated. But they are major parts of the academic body.
Today, particular fields of machine learning are more popular, such as vision, audio
or video machine learning. NLP(natural language processing), before Chat-GPT, was
much less popular and as such still has less research than other machine learning
fields. Particularly, detecting machine generated text is under researched. Though,
with the publication of Chat-GPT many more papers have been created.
4
published quickly. The papers here are already becoming outdated with the publicity
given to generated text.
With all the attention, this should continue to increase. The methodologies, also,
have changed, from a focus on traditional models to deep learning models, to a DL
variant called transformers. These are the newer, more popular ways to both generate
text and classify them.
5
With this transformer, the experiments were a “fine-tuning” of RoBERTa to classify
machine generated text as human or synthetic. Then, next, try to trick the classifier to
detect a text as human when it is fake. We found RoBERTa was very simple to trick
and the entire purpose of this thesis was to discover and propose how we would
account for how detectors fail.
6
Synthetic Text Detection: Systemic
Literature Review
Within the text analysis and processing fields, generated text attacks have been made
easier to create than ever before. To combat these attacks open sourcing models/APIs
and datasets have become a major trend to create automated detection algorithms
in defense of authenticity. For this purpose, synthetic text detection has become an
increasingly viable topic of research. This review is written for the purpose of creating
a snapshot of the state of current literature and easing the barrier to entry for future
authors. Towards that goal, we identified research trends and challenges in this field.
2.2 Introduction
Studies regarding text generation before 2017 were generally scarce and far between.
As the body of research grew for synthetic text generation, so did the research for
detectors, though lagging behind. This paper discusses current trends of research and
future viable research options with the goal of shortening the research process for
Artificial Intelligence generated text detection. The main topic of discussion is the
literature itself using the PRISMA methodology. Detection literature was reviewed
systematically and put together in detail for novel research.
7
• Shows gaps in current research for future work.
This is perhaps one of the first literature reviews on the narrow topic of generated
text detection. To the best of our knowledge there have been no systemic literature
reviews on detecting synthetic text. This study focuses on exploring the current
research literature, showing the current ecosystem behind synthetic detection and
preparing for future research.
RQ1 Which datasets are currently used in the literature to detect deep fake
models?
RQ2 What accuracy evaluation methods are there for detection effectiveness?
RQ3 What impact have recent innovations had on fake text detection?
2.3.2 Research Objectives
For this study there are 5 objectives:
8
• To explore models and datasets created to detect artificial text.
From these keywords they were assorted with AND, OR and quotation required
clauses. Certain key words like “text generation” AND “language processing” were
especially effective together at finding articles while “fake text” led to irrelevant
topics. Though some keywords were only useful for finding niche articles.
2.3.4 Article Inclusion Exclusion Criteria
A total of 1,211 articles were found using the above queries. With duplicates removed,
the remaining 1,041 articles were sifted. Many articles containing relevant keywords
and titles were not about text generation, Natural Language Processing (NLP) and
detection. Some articles were not machine centric and were about societal or human
reading of generated text though they were about generated text detection. Others
mentioned fake text as trolling, which is not the focus of this review.
In the partial review many of these papers were excluded by abstract because they
were not machine centric or were based on societal differences. Out of the 1,041
9
articles after removing duplicates, a partial review left 381 articles eligible for full
review. The following criteria was used to include or exclude these papers from this
point on:
The following is the inclusion criteria:
• Surveys on text generation are allowed The following is the exclusion criteria:
10
Figure 2.4: Publications by year
RQ1: Which datasets/models are currently used in the literature to detect generated text?
According to literature, the more training data for both human and machine generated text the
better the outcome will become. Though there are many sources for both real and fake it is good
to see the popular ones for specific domains. Below are datasets usable for training a model for
actual detection and research in the time to come:
Open-source Datasets
◦ Grover: [Link]
This is more of a collection of scripts for making a dataset oriented for news
articles. The repo includes a detection model, text generator, accuracy
evaluator and a web crawler for gathering authentic text source data.
◦ TuringBench: [Link]
11
The main website for TuringBench is more of a leaderboard about detecting
which generator is being used. The dataset is given in a zip file and whoever
gets the highest accuracy rating for detecting the generator wins.
◦ Grover: [Link]
◦ GPT-2: [Link]
face: [Link]
I mention this again, with over 5,000 models to choose from, you can see the
limitations of text generation, GPT-2 being the most popular on the platform.
◦ GLTR: [Link]
◦ Grover: [Link]
12
◦ RoFT: [Link] (gamification of human detection)
RQ2: What accuracy evaluation methods are there for detection effectiveness?
According to Li et al. [2022] a good general rule of model evaluation is testing the
mislabel or error rate. This can mean testing against a given test/validation set or
testing against an outside dataset. Using a recorded error rate you can also distinguish
how effective a detector is per generative model.
The standard way most detectors are evaluated in the literature is to stay in the
same domain and use similar datasets. The error rates are based on a narrow binary
classification and due to this there are much better error rates. You can see this
throughout the literature where there is one domain, like TweepFakeMaurizio [2021],
news oriented models Zellers et al. [2019], language based models Chen et al. [2022],
and other niche categories sticking to their own evaluated domains.
To truly test a model against more real data here are two things done in the
literature; adversarial testing Wolff [2020] and sourcing different generators Li et al.
[2022]. For adversarial testing this means accounting for a post-processing phase of
text generation whereby Greek and uncommon symbols are added along with
misspellings and surely other techniques in adversarial fashion to cause the detector
to mislabel the text as human written. Testing against this noise is a very good way to
evaluate a real-world model. A typical solution for adversarial attack is with a
preprocessing phase for the detector.
For evaluating accuracy against different sources, a good idea from Li et al. [2022]
is to record the error rates of different common and popular generator. Below is an
example of how a detector can mislabel a text,
The authors of the article found human text was most mislabeled as synthetic. This
can be used to adjust a model and evaluate its accuracy and can further be divided by
which sources the human written text originated. Though a major flaw later stated by
the author is how common it is to fine tune a generator, making labelling all
generators unfeasible.
Figure 2.6 Example evaluation based on source, ?
Lastly, how a model will be used should always be part of the evaluation. Low
resource training may be better for something like TweepFake, captioning or
13
microblog detection. Though general language detection would require more
resources and vary in effectiveness based on the text in question plus how similar the
training set had been. So the evaluation can be divided up in some cases or narrowed
in others. This way the evaluation is accurate to the model’s purpose.
RQ3: What impact have recent innovations had on fake text detection?
Across the literature the general setting of recent innovations is to combat fake text,
detecting and removing its effect on public discourse. Paid models like GPT3 are
generally superior at creating advanced generated text for fooling human detection
but is made relatively open source for research against itself and other high quality
generators.
There is a trend to open source datasets and models to protect the community
from synthetic text by making it easier to create both generation and detection
models. These motivations perhaps have had the greatest impact on future research.
Many articles, as seen in Figure 2; publications by year, have been written more
recently. Synthetic text will likely be more highly researched in the next few years
(Since writing, this has and will become more true).
Particularly niche and low resource domains are being filled in with working
models and novel solutions. Short and long form AI generated document detection
models are now more numerous with their accuracy reported in their respective
papers. These new models give us more options than the standard detector, opening
up more difficult and niche types.
4. Low resource detector optimization. Low resource training also had limited
research. TweepFakeMaurizio [2021] and fake academic paper
detectionLiyanage et al. [2022] were perhaps the most optimization related
articles. There is a gap here.
14
5. Research in other languages. Most other languages have very limited research
though some articles do existChen et al. [2022], S et al. [2022], Shamardina et
al. [2022]. Data sets exist for Chinese, Russian and other languages as well but
very few synthetic language detectors outside English.
2.7 Summary
Natural language processing is trending in good fashion with plenty of open source
projects and ideas for novelty. Automated AI text generation has grown tremendously
in the past four years and now there is a fresh need for detection. As we enter a phase
of defending trolls, bots and generated commentary recent advancements have
allowed for more of itself. In this paper, our focus is on synthetic text generation with
possible research trends and challenges.
15
A Mutation-based Text Generation for
Adversarial Machine Learning
Applications
Many natural language related applications involve text generation, created by
humans or machines. While in many of those applications machines support humans,
yet in few others, (e.g. adversarial machine learning, social bots and trolls) machines
try to impersonate humans. In this scope, we proposed and evaluated several
mutation-based text generation approaches. Unlike machinebased generated text,
mutation-based generated text needs human text samples as inputs. We showed
examples of mutation operators but this work can be extended in many aspects such
as proposing new text-based mutation operators based on the nature of the
application.
3.1 Introduction
Currently, text generation is widely used in Machine Learning (ML)-based or Artificial
Intelligence (AI)-based natural language applications such as language to language
translation, document summary, headline or abstract generation. Those applications
can be classified into different categories. In one classification, they can be divided
into short versus long text generation applications. Short text generation applications
include examples such as predicting next word or statement, image caption
generation, short language translation, and documents summarization. Long text
generation applications include long text story completion, review generation,
language translation, poetry generation, and question answering.
Large language models such as open AI chat GPT-1,2 and 3 can be used to
masquerade humans in many of those short and long text generation applications.
Unlike all previous applications where ML or AI-based text generation exists to
support humans and their applications, in online social networks (OSN) such as Twitter
and Facebook, text generation is used to fool and influence humans and public
opinions through masquerading machine-generated texts as genuinely generated by
humans. Social bots and trolls in OSNs are accounts that impersonate humans. They
are controlled by ML or AI algorithms. ML algorithms are used to auto generate text
in those bot/troll accounts.
As mentioned earlier, ML-based text generation can be used for non-malicious and
malicious applications. Our focus in this paper is in malicious applications. There are
several malicious ML-based approaches to text generation that can be observed in
literature such as:
• Classical (e.g. professor/teacher forcing recurrent neural network, RNN): It aims
to align generative behavior as closely as possible with teacher-forced behavior,
Lamb et al. [2016].
16
• Conventional (e.g. based on hidden Markov models (Jing [2002]), method of
moments (MoM), Jones [1969], Restricted Boltzmann machine (RBM), Holyoak
[1987] ).
• Cooperative training, (e.g. Yin et al. [2020], Lu et al. [2019]).
17
• Two class labels (Human/Adversarial, mutation instances are added to
adversarial instances
• Two class labels (Human/Adversarial, mutation instances are added to Human
instances
In this paper, we will propose and evaluate mutation operators within the first
category and leave investigating other categories in future papers.
The rest of the paper is organized as follows. Section 3.2 provides a summary of
related research. Our paper goals and approaches are introduced in 3.3. Section 3.4
covers the experiments we performed to evaluate our proposed mutation operators.
We have then a separate section, section 3.5 to compare with close contributions.
Finally, Section 3.6 provides some concluding remarks as well as future extensions or
directions.
18
[2018]. It is also related to AML text perturbations, Vijayaraghavan and Roy [2019],
Eger et al. [2019], Gao et al. [2018], and Li et al. [2018].
Perturbation Type Defense Example Definition
Combined Unicode ACD P.l.e.a.s.e.l.i.k.e.a.n.d. Insert a Unicode character between
s.h.a.r.e each original character.
Fake punctuation CW2V Pleas.e lik,e abd shar!e the Randomly add zero or more
v!deo punctuation marks between
characters.
Neighboring key CW2V Plwase lime and sharr the Replace character with
vvideo keyboardadjacent characters.
Random spaces CW2V Pl ease lik e and sha re th e Randomly insert zero or more
video spaces between characters. Replace
Replace Unicode UC Plea˜se lˆıke and sharˆe the characters with Unicode look-alikes
video
Space separation ACD Please l i k e and s h a r e Place spaces between characters.
UC Replace individual characters with
Tandem character obfuscation PLE/\SE LIKE /\ND
characters that together look original
SH/\RE
Transposition CW2V Please like adn sahre Swap adjacent characters Repeat
Vowel repetition and deletion CW2V Pls likee nd sharee Please or delete vowels.
ACD like and share the video Place zero-width spaces (Unicode
Zero-width space separation
character 200c) between characters
Table 3.1: Adversarial text perturbations, Bhalerao et al. [2022]
19
The random word operator is in fact a list of random words. Arbitrary words are
chosen and replaced with a random word from the same list. Limits to the number of
mutations should be added to limit the operator from completely fuzzing the string as
well. The code from our GitHub, [Link] JesseGuerrero/Mutation-Based-
Text-Detection has some written operators
Mutation Operator Example Definition
Randomization Plz shr and hate Use all below mutation
film operators
Plz sharr and like the
Misspelling words
vid Misspell a few words Delete
Please share and like a few articles, including
Deleting articles
video starting ones
Random word with Please roar and tree Replace a random word
random word video with another random word
Replace a word with its
Synonym replacement Please disseminate and
prefer the video synonym
Replace a word with its
Antonym replacement Please hide and hate
antonym
the video
Replace some a’s and e’s
Replace ”a”, ”e” Pls lik nd shar the
with epsilon & alpha
video
Table 3.2: Experimental mutation operators
20
3.4.1 Experiment methodology
We used an existing RoBERTa classifier which is meant to classify synthetic and human
generated text to test how it would classify mutated text. It is still the pre-trained
binary model, however it is being retrained to detect mutation rather than synthetic
text.
The data set used was the full COCO images data set where hand written captions
are placed for each image. A total of 5 captions are human created per image. The
captions were parsed into a re-usable format and were used to train the human
portions of the model.
Across the training, testing and validation sets there were over 700,000 human
texts used to train the model. Two models were made from this data set. The first was
based on individual captions. The second was based on these 5 captions combined
per unique image name for calling via a map.
Six operators were used as mutations for this classifier. They are; (1) replacing
synonyms, (2) antonyms, (3) random words, (4) removing articles, (5) replacing a with
alpha and e with epsilon, and lastly, (6) the most common misspellings.
The training data was duplicated for the mutation data sets. Over 700,000 texts
were used for training and the same texts were re-used for mutations. This meant for
individual captions there were over 1.4 million text instances with both mutation and
human labels. For combined captions there were 1/5 of the total instances.
The data sets were selected as they were already labelled from COCO dataset. The
training set went to training, the validation set went to validation and testing to
testing. For training the mutation label, the mutation operators were used at run-
time.
The operator was randomly chosen at run-time with a simple random function
among 6 operators. Each operator was used so the classifier can learn to detect all 6
of these mutation types in one classifier.
So far as testing is concerned, the same testing data set from COCO was re-used
with an operator manipulating a whole set. A total of 7 testing data sets were created
for each of the 6 operators and a seventh randomized data set, like the mutations at
run-time. This formed our metrics of how accurately the model can correctly label
each instance of the mutation data sets as mutations and how accurately the model
can label human text. In a total 8 operators, 1 human and a seventh randomizing the
first six mutations.
∼Randomized
01.00%(1000 samples)
21
Replace Alpha, Epsilon ∼01.01%(1000 samples)
Misspelling words ∼00.00%(1000 samples)
Delete articles ∼01.60%(1000 samples)
Synonym replacement ∼00.00%(1000 samples)
Replace random word ∼07.79%(1000 samples)
Antonym replacement ∼09.89%(1000 samples)
Table 3.3: Preliminary Results
As we can see from 1000 samples the accuracy is quite poor when modifying the
text. The original detector without mutations had a recall of over 97% detection of
synthetic and human text in-distribution. For our research outside of the paper we
have an out of distribution pure human text data set as 88% accurate as the 1st row
in the table.
This means the detector is quite good at classifying human text out of distribution
and is even good at detecting in-distribution synthetic text. The model does those
things above human distinction which in the past was around 54% accurate. Our issue
from our modeling is mutation from which the model does not perform well.
∼99.83%(2490)
Randomized
Replace alpha, epsilon ∼99.95%(2490)
Misspelling words Delete articles
59.87%(2490)
Synonym replacement ∼99.91%(2490)
Replace random word ∼100%(2490)
Antonym replacement ∼99.03%(2490)
Table 3.4: Individual captions, short language modeling
*For the live thesis defense a different model was trained as the old was lost. The
live results differ by 10% for overall accuracy.
The most inaccurate operator overall was always the ”delete articles operator”.
This can be due to some semantic issues or just the difficulty of detecting what is not
there rather than what is there to a RoBERTa classifier. For this first run, the other
operators ranged from 59% accurate to 100% accurate, with human detection being
71%.
22
For the second run the captions for text chunks were combined per image. This
meant all 5 captions were now one text and were fed to the training model. This
means 1/5 th of the instances but more per text. The results were a slightly lower;
total overall detector accuracy of 88% with 2490 texts being tested.
The epochs took about 2 hours each this way as well. A total of 10 epochs or 20
hours were used to finish the model training. Same as before, the delete articles
operator was the weakest mutator and was in fact even weaker the second time.
Besides the weakest operators, the other operators were all above 95%, much
better than the original neural network. If we were to remove the lowest performing
operator we would in fact have 95% accuracy for the 1st run and 97%
∼98.96%(2490)
Randomized
Replace alpha, epsilon ∼99.92%(2490)
Misspelling words ∼99.80%(2490)
Delete articles ∼25.42%(2490)
Synonym replacement ∼99.76%(2490)
Replace random word ∼98.43%(2490)
Antonym replacement ∼92.37%(2490)
Table 3.5: Combined captions, longer language modeling
accuracy for the 2nd. This means in reality the combined text may be the better
approach to train.
23
3.5 A Comparison Study
3.5.1 Machine Text Generation
Machine text generation is a field of study in Natural Language Processing (NLP) that
combines computational linguistics and artificial intelligence that has progressed
significantly in recent years. Neural network approaches are widely used for this task
and keep dominant in the field. The state-of-the-art methods may include
Transformers Vaswani et al. [2017], BERT Devlin et al. [2018], GPT3 Brown et al.
[2020], RoBERTa Liu et al. [2019], etc. The models are trained on a large amount of
text data. For example, the GPT-2 model was trained on text scrapped from eight
million web pages Radford et al. [2019], and is able to generate human-like text. Due
to the high text generation performance, such methods are very popular on tasks,
such as image caption generation, text summarization, machine translation, moving
script-writing, and poetry composition. However, the output of such methods is often
open-ended.
Different papers show promising results regarding transfomers, ensemble learn
ing and RNNs. Works like Li et al. [2022], Tourille et al. [2022] and Wolff [2020] include
the ability to detect GPT models very accurately. In Li et al. [2022] for example, the
ensemble model is able to detect which model is being used to generate text with
66% accuracy, not just what is synthetic and what is real. Author attribution is difficult,
so this is a relatively high number.
In Tourille et al. [2022] transformers are used to detect generated tweets from
different models with an above 90% accuracy. Similar results were found in NajeeUllah
et al. [2022] with an above 95% accuracy. It is common to find papers based on
transformers with such great results.
There are perhaps a few dozen papers all showing amazing accuracy for
transformers. However there are much less on adversarial attacks and only 1 or 2 on
mutating generated texts for detection. This is why this paper was written, to
contribute to this topic which is scarce in research.
Through this work, we propose a mutation-based text generation method that can
be distinguished from the existing text generation method fundamentally. Unlike the
neural network based methods, the mutation-based method generates output based
on the given text under a given condition. The text is generated in a tightly controlled
environment, and the output is closed-ended. The wellcontrolled environment makes
the output of the mutation-based method suitable for serious security test tasks, such
as machine learning model vulnerability tests and SQL injection defense and
detection.
Given a text corpus (e.g., a sentence or a paragraph), T , which contains an ordered
set of words, W = {w1,w2,...,wn}, and an ordered set of punctuation, P = {p1,p2,...,pm},
a mutation operator, µ(·) is used to generate the mutationbased text. For instance,
given a character-level mutation operator, µc(·):
W′ = µc(wi,ρ,σ), (3.1)
24
where W′ is the output of µc(·), which replaces the letter ρ in wi (wi ∈ W) with a
mutation σ. Then, the final output of T is T ′ = ⟨W′,P⟩. For instance, assume
T = "Text generation is interesting!", W = {Text,generation,is,interesting}, and P = {!}, wi =
generation, and a character-level mutation operator µu(·), where ρ = a and σ = α (the Greek letter
alpha). Then, the output text corpus, T ′ is generated as:
T ′ = ⟨W′,P⟩ =
⟨µc(wi,ρ,σ),P⟩
(3.2)
= ⟨µc(generation,a,α),P⟩
= Text generαtion is interesting!
3.6 Conclusions
Automatic text generation techniques are adopted into various domains, from
question-answering to AI-driven education. Due to the progress of neural network
(NN) techniques, NN-based approaches dominate the field. Though advanced
techniques may be applied to control text generation direction, the text is still
generated in a widely open-ended fashion. For instance, a NN-based approach can
generate a greeting message to greet a specific person. However, it is hard to control
the exact wording used in the message. Such open-ended text generation might work
fine for content generation tasks. However, due to lack of precise control, using open-
25
ended text generation methods to systematically evaluate flaws in language analysis
models may be non-trivial.
Unlike the existing language models, our proposed mutation-based text-generation
framework provides a tightly controlled environment for text generation that extends text-
generation techniques to the field of cyber security (i.e., flaws evaluation for language
analysis models). The output of our framework can be used to systematically evaluate any
machine learning models or software systems that use a sequence of text as input, such
as SQL injection detection Hlaing and Khaing [2020] and software debugging Zhao et al.
[2022]. Researchers may also design their own mutation operators under our framework.
We demonstrated the proposed text-generation framework using the RoBERTa
based detector that is pre-trained for separating human-written text from synthetic
ones. Our experiments showed that the RoBERTa-based detector has a significant flaw.
As a detection method, it is extremely vulnerable to simple adversarial attacks, such as
replacing the English letter ”a” with the Greek letter ”α” or removing the articles—a,
an, the—from a sentence. We also demonstrated that simply including the adversarial
samples (i.e., the mutation texts) in the finetuning stage of the classifier would
significantly improve the model robustness on such types of attack. However, we
believe that this issue should be better addressed on the feature level since any changes
at the text level will lead to changes in the tokenization stage that will eventually lead
to a different embedding vector being fed into the classification network. Thus, one
future direction of this work is reducing the distance between the original and mutation
samples in the feature space. Some potentially useful methods might include using
contrastive learning and siamese network Koch et al. [2015], Liang et al. [2021] as well
as dynamic feature alignment Zhang et al. [2022], Dong et al. [2020].
Besides improving the robustness of the RoBERTa-based detector, we plan to
continue to work on the development of the proposed mutation-based text
generation framework. Currently, the framework only works in a two-step testing
scenario. Users need to use the framework to generate the testing cases and feed the
testing cases into the downstream model in separate steps. We plan to release a
library that can be directly imported into any downstream applications. Tools for easily
creating and editing mutation operators will also be created. In addition, a graphical
user interface may also be developed.
In conclusion, we propose a general-purpose, mutation-based text generation
framework that produces close-ended, precise text. The output of our framework can
be used in various downstream applications that take text sequences as input,
providing a systematic way to evaluate the robustness of such models. We believe the
proposed framework offers a new direction to systematically evaluate language
models that will be very useful to those who are seeking insightful analysis of such
models.
26
28
3.6.1 Summary
In chapter 2 we cover how fake text detection is open to research. We also stated,
within Computer Science some fields also are saturated, such as Computer Vision and
Computer Networks. However, for machine centric synthetic text detection, as in
machine on machine generation and detection, there are plenty of topics to claim as
contribution. In chapter 2 you can see what research avenues are most available.
In chapter 3 we use one of these research avenues found in the Systemic
Literature Review. Particularly, adversarial attacks on detectors using mutations. As
said in the chapter the normal accuracy of these transformer detectors has a 97%
recall rate, quite high. By simply adding mutations the accuracy moves below 10%.
This is the question we seek to answer in chapter 3, ”can we defend against
mutations?” The answer is yes, we can. The experiment successfully showed we could
learn 6 types of mutations and classify them as mutation. Future work from chapter
3 would then become two parts, finding more operators that are either more human
or more AI-oriented, and also, combining the mutation label with synthetic to detect
fake text that has been mutated under the right label.
3.6.2 Conclusion
The greatest barrier to entry for new academics will be learning to create models from
neural networks and understanding the models themselves outside a black box. There
are plenty of viable options for academia, including fake foreign language detection,
low resource learning, detector adversarial attacks, domain specific attacks such as
fake news and student essays and generalization. Because of this, the trend of today
is rapid growth of all NLP fields at an increasing rate.
As for adversarial attacks, the COCO dataset was a great way to test how to
overcome an attack on a detector based on mutation. The original under 10%
accuracy of an attacked detector was boosted to 85%+ by having a heuristic as to what
that mutation would be. This can be applied to any detector for any number of
mutation operators.
This would be particularly helpful for students who write essays with GPT-3.5 and
whom edit the essay to trick their professor. The future research in question from that
chapter will be coming up with heuristics which accurately catch these deep fake
texts.
Bibliography
David Ifeoluwa Adelani, Haotian Mai, Fuming Fang, Huy H. Nguyen, Junichi Yamagishi,
and Isao Echizen. Generating sentiment-preserving fake online reviews using neural
language models and their human- and machine-based detection,
2020. URL [Link]
Izzat Alsmadi, Nura Aljaafari, Mahmoud Nazzal, Shadan Alhamed, Ahmad H.
Sawalmeh, Conrado P. Vizcarra, Abdallah Khreishah, Muhammad Anan, Abdulelah
Algosaibi, Mohammed Abdulaziz Al-Naeem, Adel Aldalbahi, and Abdulaziz Al-
Humam. Adversarial machine learning in text processing: A literature survey. IEEE
Battista Biggio, Blaine Nelson, and Pavel Laskov. Poisoning attacks against support
vector machines. arXiv preprint arXiv:1206.6389, 2012.
J´er´emie Bogaert, Marie-Catherine de Marneffe, Antonin Descampe, and
FrancoisXavier Standaert. Automatic and manual detection of generated news:
Case study, limitations and challenges. pages 18–26. ACM, 6 2022. ISBN
9781450392426. doi: 10.1145/3512732.3533589. URL [Link]
doi/10.1145/3512732.3533589.
28
Sergio Bravo-Santos, Esther Guerra, and Juan de Lara. Testing chatbots with
charm. In International Conference on the Quality of Information and
Communications Technology, pages 426–438. Springer, 2020.
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan,
Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell,
et al. Language models are few-shot learners. Advances in neural information
processing systems, 33:1877–1901, 2020.
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. Evaluation of text generation: A
survey. 6 2020. doi: 10.48550/arxiv.2006.14799. URL https:
//[Link]/abs/2006.14799v2.
Tong Che, Yanran Li, Ruixiang Zhang, R Devon Hjelm, Wenjie Li, Yangqiu Song, and
Yoshua Bengio. Maximum-likelihood augmented discrete generative adversarial
networks. arXiv preprint arXiv:1702.07983, 2017.
Liqun Chen, Shuyang Dai, Chenyang Tao, Haichao Zhang, Zhe Gan, Dinghan Shen,
Yizhe Zhang, Guoyin Wang, Ruiyi Zhang, and Lawrence Carin. Adversarial text
generation via feature-mover’s distance. Advances in Neural Information
Processing Systems, 31, 2018.
X Chen, P Jin, S Jing, C Xie 2022 IEEE 10th Joint International, and undefined 2022.
Automatic detection of chinese generated essayss based on pretrained bert.
[Link], 2022. URL [Link]
abstract/document/9836571/.
Alesia Chernikova and Alina Oprea. Fence: Feasible evasion attacks on neural
networks in constrained environments. ACM Transactions on Privacy and Security,
25(4):1–34, 2022.
Ayesha Priyambada Das, Ajit Kumar Nayak, and Mamata Nayak. A survey on machine
learning based text categorization. [Link], 2018. doi: 10.9790/0661-
2002035156. URL [Link]
56747916/[Link].
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pretraining
of deep bidirectional transformers for language understanding. arXiv preprint
arXiv:1810.04805, 2018.
Jiahua Dong, Yang Cong, Gan Sun, Yuyang Liu, and Xiaowei Xu. Cscl: Critical semantic-
consistent learning for unsupervised domain adaptation. In European Conference
on Computer Vision, pages 745–762. Springer, 2020.
Liam Dugan, Daphne Ippolito, Arun Kirubarajan, and Chris Callison-Burch. Roft: A tool
for evaluating human detection of machine-generated text. pages
189–196, 10 2020. doi: 10.48550/arxiv.2010.03070. URL [Link]
abs/2010.03070v1.
Steffen Eger, Go¨zde Gu¨l S¸ahin, Andreas Ru¨ckl´e, Ji-Ung Lee, Claudia Schulz,
Mohsen Mesgar, Krishnkant Swarnkar, Edwin Simpson, and Iryna Gurevych. Text
processing like humans do: Visually attacking and shielding nlp systems. arXiv
preprint arXiv:1903.11508, 2019.
Noureen Fatima, Ali Shariq Imran, Zenun Kastrati, Sher Muhammad Daudpota, and
Abdullah Soomro. A systematic literature review on text generation using deep
neural network models. IEEE Access, 10:53490–53503, 5 2022. doi: 10.
1109/ACCESS.2022.3174108.
Matthias Gall´e, Jos Rozen, Germ´an Kruszewski, and Hady Elsahar. Unsupervised and
distributional detection of machine-generated text. [Link], 11 2021. URL
[Link]
Margherita Gambini, Tiziano Fagni, Fabrizio Falchi, and Maurizio Tesconi. On pushing
deepfake tweet detection capabilities to the limits. pages 154–163. ACM, 6 2022.
ISBN 9781450391917. doi: 10.1145/3501247.3531560. URL
[Link]
Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. Black-box generation of
adversarial text sequences to evade deep learning classifiers. In 2018 IEEE Security
and Privacy Workshops (SPW), pages 50–56. IEEE, 2018.
Zar Chi Su Su Hlaing and Myo Khaing. A detection and prevention technique
30
on sql injection attacks. In 2020 IEEE Conference on Computer Applications (ICCA),
pages 1–6. IEEE, 2020.
Keith J Holyoak. Parallel distributed processing: explorations in the microstructure of
cognition. Science, 236:992–997, 1987.
Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. Automatic
detection of generated text is easiest when humans are fooled. pages 1808–1822,
11 2019. doi: 10.48550/arxiv.1911.00650. URL [Link]
org/abs/1911.00650v2.
Touseef Iqbal and Shaima Qureshi. The survey: Text generation models in deep
learning. Journal of King Saud University - Computer and Information Sciences,
34:2515–2528, 6 2022. ISSN 1319-1578. doi: 10.1016/[Link].2020. 04.001.
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. Adversarial example
generation with syntactically controlled paraphrase networks. arXiv preprint
arXiv:1804.06059, 2018.
Matthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu, Cristina NitaRotaru, and Bo
Li. Manipulating machine learning: Poisoning attacks and countermeasures for
regression learning. In 2018 IEEE Symposium on Security and Privacy (SP), pages
19–35. IEEE, 2018.
Ganesh Jawahar, Muhammad Abdul-Mageed, V.S. Laks Lakshmanan, M AbdulMageed
arXiv preprint arXiv ..., and undefined 2020. Automatic detection of machine
generated text: A critical survey. [Link], pages 2296–2309, 1 2020. doi:
10.18653/V1/[Link]-MAIN.208. URL [Link]
2011.01314[Link]
31
Alex M Lamb, Anirudh Goyal ALIAS PARTH GOYAL, Ying Zhang, Saizheng Zhang, Aaron
C Courville, and Yoshua Bengio. Professor forcing: A new algorithm for training
recurrent networks. Advances in neural information processing systems, 29, 2016.
Brandon Laughlin, Christopher Collins, Karthik Sankaranarayanan, and Khalil El-Khatib.
A visual analytics framework for adversarial text generation. In 2019 IEEE
Symposium on Visualization for Cyber Security (VizSec), pages 1–10. IEEE, 2019.
Thomas Lavergne, Tanguy Urvoy, Fran¸cois Yvon, T Lavergne, A T Urvoy, and´ F Yvon.
Filtering artificial texts with statistical machine learning techniques. Springer,
45:25–43, 3 2011. doi: 10.1007/s10579-009-9113-0. URL https:
//[Link]/article/10.1007/s10579-009-9113-0.
Thai Le, Suhang Wang, and Dongwon Lee. Malcom: Generating malicious comments
to attack neural fake news detection models. 2020 IEEE International Conference
on Data Mining (ICDM), 2020-Novem:282–291, 8 2020. URL
[Link]
Bin Li, Yixuan Weng, Qiya Song, and Hanjun Deng. Artificial text detection with
multiple training strategies. [Link], 2022. doi: 10.28995/ 2075-7182-2022-
20-375-381. URL [Link] [Link].
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. Textbugger: Generating
adversarial text against real-world applications. arXiv preprint arXiv:1812.05271,
2018.
Junyi Li, Tianyi Tang, Wayne Xin Zhao, and Ji Rong Wen. Pretrained language models
for text generation: A survey. IJCAI International Joint Conference on Artificial
Intelligence, pages 4492–4499, 1 2021. ISSN 10450823. doi: 10.24963/
ijcai.2021/612. URL [Link]
[Link]/abs/2201.05273v4.
Gongbo Liang, Connor Greenwell, Yu Zhang, Xin Xing, Xiaoqin Wang, Ramakanth
Kavuluru, and Nathan Jacobs. Contrastive cross-modal pre-training: A general
strategy for small sample medical imaging. IEEE Journal of Biomedical and Health
Informatics, 26(4):1640–1649, 2021.
Kevin Lin, Dianqi Li, Xiaodong He, Zhengyou Zhang, and Ming-Ting Sun. Adversarial
ranking for language generation. Advances in neural information processing
systems, 30, 2017.
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy,
Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized
bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
32
Zhengyuan Liu and Nancy Chen. Entity-based de-noising modeling for controllable
dialogue summarization. In Proceedings of the 23rd Annual Meeting of the Special
Interest Group on Discourse and Dialogue, pages 407–418, 2022.
Sidi Lu, Yaoming Zhu, Weinan Zhang, Jun Wang, and Yong Yu. Neural text generation:
Past, present and beyond. 3 2018. URL [Link]
1803.07133.
Sidi Lu, Lantao Yu, Siyuan Feng, Yaoming Zhu, and Weinan Zhang. Cot: Cooperative
training for generative modeling of discrete data. In International Conference on
Machine Learning, pages 4164–4172. PMLR, 2019.
Narek Maloyan, Lomonosov Msu, Bulat Nutfullin, and Eugene Ilyushin. Dialog22 ruatd
generated text detection. 6 2022. URL [Link]
2206.08029.
AS Nayak. Deepspot: spotting fake reviews with sentiment analysis and text
generation. 2019. URL [Link]
view/pdfCoverPage?instCode=01CALS_USL&filePid=13232648180001671&
download=true.
Blaine Nelson, Marco Barreno, Fuching Jack Chi, Anthony D Joseph, Benjamin IP
Rubinstein, Udam Saini, Charles Sutton, J Doug Tygar, and Kai Xia. Exploiting
machine learning to subvert your spam filter. LEET, 8(1):9, 2008.
Andrew Newell, Rahul Potharaju, Luojie Xiang, and Cristina Nita-Rotaru. On the
practicality of integrity attacks on document-level sentiment analysis. In
Proceedings of the 2014 Workshop on Artificial Intelligent and Security Workshop,
pages 83–93, 2014.
James Newsome, Brad Karp, and Dawn Song. Paragraph: Thwarting signature learning
by training maliciously. In International Workshop on Recent Advances in Intrusion
Detection, pages 81–105. Springer, 2006.
33
[Link]
Roberto Perdisci, David Dagon, Wenke Lee, Prahlad Fogla, and Monirul Sharif.
Misleading worm signature generators using deliberate noise injection. In 2006
IEEE Symposium on Security and Privacy (S&P’06), pages 15–pp. IEEE, 2006.
Saad Ahmed Qazi, Hina Kirn, Muhammad Anwar, Ashina Sadiq, Hafiz M Zeeshan,
Imran Mehmood, and Rizwan Aslam Butt. Deepfake tweets detection using deep
learning algorithms. [Link], 2022. doi: 10.3390/ engproc2022020002. URL
[Link]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et
al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9,
2019.
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. Semantically equivalent
adversarial rules for debugging nlp models. In Annual Meeting of the Association
for Computational Linguistics (ACL), 2018.
Benjamin IP Rubinstein, Blaine Nelson, Ling Huang, Anthony D Joseph, Shinghon Lau,
Satish Rao, Nina Taft, and J Doug Tygar. Antidote: understanding and defending
against poisoning of anomaly detectors. In Proceedings of the 9th ACM SIGCOMM
Conference on Internet Measurement, pages 1–14, 2009.
Paolo Russu, Ambra Demontis, Battista Biggio, Giorgio Fumera, and Fabio Roli. Secure
kernel machines against evasion attacks. In Proceedings of the 2016 ACM workshop
on artificial intelligence and security, pages 59–69, 2016.
Sina Mahdipour Saravani, Indrakshi Indrajit Ray, and Indrakshi Indrajit Ray. Automated
identification of social media bots using deepfake text detec-
tion. Lecture Notes in Computer Science (including subseries Lecture Notes in
Artificial Intelligence and Lecture Notes in Bioinformatics), 13146 LNCS: 111–123,
2021. ISSN 16113349. doi: 10.1007/978-3-030-92571-0 7. URL
[Link] Ali Sayghe,
Junbo Zhao, and Charalambos Konstantinou. Evasion attacks with adversarial deep
learning against power system state estimation. In 2020 IEEE Power & Energy Society
General Meeting (PESGM), pages 1–5. IEEE, 2020.
T Schuster, R Schuster, DJ Shah Computational ..., and undefined 2020. The limitations
of stylometry for detecting machine-generated fake news. [Link], 2020.
34
doi: 10.1162/COLI. URL [Link] article-
abstract/46/2/499/93369.
Dingmeng Shi, Zhaocheng Ge, and Tengfei Zhao. Word-level textual adversarial
attacking based on genetic algorithm. In Third International Conference on
Computer Communication and Network Security (CCNS 2022), volume 12453, pages
272–276. SPIE, 2022.
Rui Shu, Tianpei Xia, Laurie Williams, and Tim Menzies. Omni: automated ensemble
with unexpected models against adversarial evasion attack. Empirical Software
Engineering, 27(1):1–32, 2022.
Reuben Tan, Bryan A. Plummer, and Kate Saenko. Detecting cross-modal inconsistency
to defend against neural fake news. 9 2020. doi: 10.18653/v1/2020. emnlp-
main.163. URL [Link]
org/10.18653/v1/[Link]-main.163.
Chen Tang, Frank Guerin, Yucheng Li, C Lin arXiv preprint arXiv:2203.03047, undefined
2022, and Chenghua Lin. Recent advances in neural text generation: A task-agnostic
survey. [Link], 3 2022. URL [Link]
03047[Link]
Julien Tourille, Babacar Sow, Adrian Popescu, A Popescu of the 1st International
Workshop on ..., undefined 2022, and Adrian Popescu. Automatic detection of bot-
generated tweets. [Link], pages 44–51, 6 2022. doi: 10.1145/3512732.
3533584. URL [Link]
35
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N
Gomez, L ukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in
neural information processing systems, 30, 2017.
abs/2002.11768.
Congying Xia, Chenwei Zhang, Hoang Nguyen, Jiawei Zhang, and Philip Yu. Cg-bert:
Conditional text generation with bert for generalized few-shot intent detection.
[Link], 4 2020. URL [Link]
Haiyan Yin, Dingcheng Li, Xu Li, and Ping Li. Meta-cotgan: A meta cooperative training
paradigm for improving adversarial text generation. In Proceedings of the AAAI
Conference on Artificial Intelligence, volume 34, pages 9466–9473, 2020.
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska
Roesner, and Yejin Choi. Defending against neural fake news.
Advances in Neural Information Processing Systems, 32, 5 2019. ISSN 10495258.
doi: 10.48550/arxiv.1905.12616. URL [Link]
1905.12616v3.
Fuyong Zhang, Yi Wang, Shigang Liu, and Hua Wang. Decision-based evasion attacks
on tree ensemble classifiers. World Wide Web, 23(5):2957–2977, 2020.
Yu Zhang, Gongbo Liang, and Nathan Jacobs. Dynamic feature alignment for semi-
supervised domain adaptation. British Machine Vision Conference (BMVC), 2022.
Yu Zhao, Ting Su, Yang Liu, Wei Zheng, Xiaoxue Wu, Ramakanth Kavuluru, William GJ
Halfond, and Tingting Yu. Recdroid+: Automated end-to-end crash reproduction
36
from bug reports for android apps. ACM Transactions on Software Engineering and
Methodology (TOSEM), 31(3):1–33, 2022.
Wanjun Zhong, Duyu Tang, Zenan Xu, Ruize Wang, Nan Duan, Ming Zhou, Jiahai Wang,
and Jian Yin. Neural deepfake detection with factual structure of text. [Link], 10
2020. URL [Link]
37