Dissertation on
“Towards Equitable Text Generation:
Bias Detection and Mitigation”
Submitted in partial fulfilment of the requirements for the award of the degree
of
Bachelor of Technology
In
CSE(AI&ML)
UE22AM320A – Capstone Project Phase - II
Submitted by:
Name - Kavya V SRN1- PES1UG22AM083
Name - Samprith Jagtap D SRN2- PES1UG22AM145
Name- Suraj B M SRN3- PES1UG22AM172
Name- T Jaiwanth SRN4- PES1UG22AM173
Under the guidance of
Prof. Bhaskar Jyoti Das
Adjunct Professor
PES University
January - May 2025
DEPARTMENT OF CSE(AI&ML)
FACULTY OF ENGINEERING
PES UNIVERSITY
(Established under Karnataka Act No. 16 of 2013)
100 feet Ring road, BSK 3rd stage, Hosakerehalli, Bengaluru – 560085
PES UNIVERSITY
(Established under Karnataka Act No. 16 of 2013)
100ft Ring Road, Bengaluru – 560 085, Karnataka, India
FACULTY OF ENGINEERING
CERTIFICATE
This is to certify that the dissertation entitled
‘Towards Equitable Text Generation:
Bias Detection and Mitigation’
is a bonafide work carried out by
Name- Kavya V PES1UG22AM083
Name- Samprith Jagtap D PES1UG22AM145
Name- Suraj B M PES1UG22AM172
Name- T Jaiwanth PES1UG22AM173
In partial fulfilment for the completion of Sixth-semester Capstone Project Phase - II
(UE22AM320A) in the Program of Study -Bachelor of Technology in CSE(AI&ML) under
rules and regulations of PES University, Bengaluru during the period January 2025 – May
2025. It is certified that all corrections/suggestions indicated for internal assessment have
been incorporated in the report. The dissertation has been approved as it satisfies the
6th-semester academic requirements in respect of project work.
Signature Signature Signature
Prof. Bhaskar Jyoti Das Dr. Jayashree R Dean, Engineering and
Adjunct Professor Chairperson Technology
External Viva
Name of the Examiners Signature with Date
1. __________________________ __________________________
2. __________________________ __________________________
DECLARATION
We hereby declare that the Capstone Project Phase - II entitled “Towards Equitable
Text Generation: Bias Detection and Mitigation” has been carried out by us under
the guidance of Prof. Bhaskar Jyoti Das, Professor and submitted in partial
fulfilment of the course requirements for the award of the degree of Bachelor of
Technology in CSE(AI&ML) of PES University, Bengaluru during the academic
semester January – May 2025. The matter embodied in this report has not been
submitted to any other university or institution for the award of any degree.
SRN1- PES1UG22AM083 Name - Kavya V
SRN2- PES1UG22AM145 Name - Samprith Jagtap D
SRN3- PES1UG22AM172 Name- Suraj B M
SRN4- PES1UG22AM173 Name- T Jaiwanth
ACKNOWLEDGEMENT
I would like to express my gratitude to Prof. Bhaskar Jyoti Das, Department of CSE(AI&ML),
PES University, for his/ her continuous guidance, assistance, and encouragement throughout the
development of UE22CS320A - Capstone Project Phase – II.
I am grateful to Capstone Project Coordinator, Prof. Shwetha K N for organizing, managing, and
helping with the entire process.
I take this opportunity to thank Dr. Jayashree R, Professor & Chairperson, Department of
CSE(AI&ML), PES University, for all the knowledge and support I have received from the
department. I would like to thank Dean, Engineering and Technology, PES University for his help.
I am deeply grateful to Prof. Jawahar Doreswamy, Chancellor, PES University, Dr. Suryaprasad J,
Vice-Chancellor, PES University, for providing me with various opportunities and enlightenment
every step of the way. Finally, Phase - II of the project could not have been completed without the
continual support and encouragement I have received from my family and friends.
ABSTRACT
Natural Language Generation (NLG) models have made impressive strides in producing text that
mimics human language. However, these models often inherit and amplify societal
biases—especially those related to gender, profession, and race—that are embedded in their training
data. As a result, AI-generated content can unintentionally reinforce harmful stereotypes, raising
serious ethical and social concerns. To address this, our project introduces CEDAR (Context-Aware
Embedding-Based Dynamic Adversarial Reinforcement learning), a novel, multi-stage framework
designed to detect and mitigate biases in NLG systems. CEDAR combines three key components: a
BERT-based bias detection module that identifies contextual bias, adversarial training that
dynamically reduces bias during model learning, and reinforcement learning with fairness-aware
reward functions to guide the generation of fluent and unbiased text. Our framework is designed to
strike a balance between ethical responsibility and linguistic performance—ensuring that generated
outputs remain coherent and semantically relevant while minimizing both explicit and implicit
biases. This work aims to advance the development of more responsible, inclusive, and trustworthy
AI systems.
TABLE OF CONTENTS
Chapter No. Title Page No.
1. INTRODUCTION 01
2. PROBLEM DEFINITION 02
3. EXTENDED LITERATURE SURVEY 03
3.1 “Towards Robust NLG Bias Evaluation with Syntactically-diverse
Prompts”
3.2 “Societal Biases in Language Generation: Progress and Challenges”
3.3 “The Woman Worked as a Babysitter: On Biases in Language
Generation”
3.4 “This is a Problem, Don’t You Agree? Framing and Bias in Human
Evaluation for Natural Language Generation”
3.5 “Adhering, Steering, and Queering: Treatment of Gender in Natural
Language Generation”
3.6 “ Bias mitigation for large language models using adversarial
learning”
3.7 “NLPositionality: Characterizing design biases of datasets and
models”
3.8 “Undesirable biases in NLP: Addressing challenges of
measurement”
3.9 “Data augmentation techniques in natural language processing”
4. HIGH LEVEL DESIGN 05
5. DATA PRE-PROCESSING 06
6. IMPLEMENTATION 07
7. CONCLUSION OF CAPSTONE PROJECT PHASE – II 12
8. PLAN OF WORK FOR CAPSTONE PROJECT PHASE - III 13
LIST OF FIGURES
Figure No. Title Page
No.
Figure- 1 High Level Design 5
Figure -2 Data Pre-Processing 6
Figure - 3 Implementation (5.a) 7
Figure - 4 Implementation (5.b) 8
CHAPTER 1
INTRODUCTION
1.1
We implemented our bias-free NLG model using a layered approach focused on fairness, context,
and text quality. First, we built a synthetic dataset of profession-based prompts labeled for bias
presence, simulating real-world scenarios. A BERT-based classifier was trained to detect both
explicit and subtle biases by understanding contextual embeddings. Then, we introduced adversarial
debiasing using a discriminator network that flagged biased outputs from the T5 decoder. When bias
was detected, a penalty was backpropagated to guide the model toward fairer language. To further
improve fairness, we fine-tuned the T5 model using reinforcement learning with a custom reward
function combining fluency and fairness. By carefully tuning the penalty weights, we preserved text
coherence while reducing biased outputs. The model was trained over multiple epochs and evaluated
using a fixed set of prompts. We observed significant bias reduction compared to the original model,
with minimal impact on output fluency. The entire pipeline, including model saving and reloading,
ensures repeatability and integration readiness. Overall, our system combines context-aware
detection, adversarial feedback, and reinforcement learning to produce fair, fluent, and
responsible text.
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025 1
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
2.
PROBLEM DEFINITION
NLG models create biased content because of unbalanced data and model limitations. The biases
perpetuate harmful stereotypes and lower the credibility of AI.
The task is to:
Identify and minimize implicit and explicit biases in NLG models without compromising text quality
and contextual appropriateness.
This demands to tackle:
● Subtle bias detection,
● Fairness without overcorrection,
● High computational cost,
● Subjectivity in defining fairness.
Our method introduces a strong, scalable framework to provide fair and ethical AI text generation.
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
3. EXTENDED LITERATURE SURVEY
1. Aggarwal, A., Sun, J. and Peng, N., 2022. Towards robust NLG bias evaluation with
syntactically-diverse prompts. arXiv preprint arXiv:2212.01700.
2. Sheng, E., Chang, K.W., Natarajan, P. and Peng, N., 2021. Societal biases in language
generation: Progress and challenges. arXiv preprint arXiv:2105.04054.
3. Sheng, E., Chang, K.W., Natarajan, P. and Peng, N., 2019. The woman worked as a babysitter:
On biases in language generation. arXiv preprint arXiv:1909.01326.
4. Schoch, S., Yang, D. and Ji, Y., 2020, December. “This is a Problem, Don’t You Agree?”
Framing and Bias in Human Evaluation for Natural Language Generation. In Proceedings of the 1st
Workshop on Evaluating NLG Evaluation (pp. 10-16).
5. Strengers, Y., Qu, L., Xu, Q. and Knibbe, J., 2020, April. Adhering, steering, and queering:
Treatment of gender in natural language generation. In Proceedings of the 2020 CHI Conference on
Human Factors in Computing Systems (pp. 1-14).
6. Dhamala, J., Sun, T., Kumar, V., Krishna, S., Pruksachatkun, Y., Chang, K.W. and
Gupta, R., 2021, March. Bold: Dataset and metrics for measuring biases in open-ended language
generation. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency
(pp. 862-872).
7. Wei, J.T.Z. and Jia, R., 2021. The statistical advantage of automatic NLG metrics at the
system level. arXiv preprint arXiv:2105.12437.
8. Das, M. and Balke, W.T., 2022, October. Quantifying bias from decoding techniques in
natural language generation. In Proceedings of the 29th International Conference on Computational
Linguistics (pp. 1311-1323).
9. Zhou, K., Blodgett, S.L., Trischler, A., Daumé III, H., Suleman, K. and Olteanu, A., 2022.
Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications. arXiv
preprint arXiv:2205.06828.
10. Sahoo, N., Gupta, H. and Bhattacharyya, P., 2022. Detecting unintended social bias in toxic
language datasets. arXiv preprint arXiv:2210.11762.
11. Dhamala, J., Sun, T., Kumar, V., Krishna, S., Pruksachatkun, Y., Chang, K.W. and Gupta, R.,
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
2021, March. Bold: Dataset and metrics for measuring biases in open-ended language generation. In
Proceedings of the 2021 ACM conference on fairness, accountability, and transparency (pp.
862-872).
12. Wei, J.T.Z. and Jia, R., 2021. The statistical advantage of automatic NLG metrics at the
system level. arXiv preprint arXiv:2105.12437.
13. Das, M. and Balke, W.T., 2022, October. Quantifying bias from decoding techniques in
natural language generation. In Proceedings of the 29th International Conference on Computational
Linguistics (pp. 1311-1323).
14. Zhou, K., Blodgett, S.L., Trischler, A., Daumé III, H., Suleman, K. and Olteanu, A., 2022.
Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications. arXiv
preprint arXiv:2205.06828.
15. Sahoo, N., Gupta, H. and Bhattacharyya, P., 2022. Detecting unintended social bias in toxic
language datasets. arXiv preprint arXiv:2210.11762.
16 Guo, Y., Guo, M., Su, J., Yang, Z., Zhu, M., Li, H., Qiu, M. and Liu, S.S., 2024. Bias in large
language models: Origin, evaluation, and mitigation. arXiv preprint arXiv:2411.10915.
17. Pellicer, L.F.A.O., Ferreira, T.M. and Costa, A.H.R., 2023. Data augmentation techniques in
natural language processing. Applied Soft Computing, 132, p.109803.
18. Van der Wal, O., Bachmann, D., Leidinger, A., van Maanen, L., Zuidema, W. and Schulz, K.,
2024. Undesirable biases in NLP: Addressing challenges of measurement. Journal of Artificial
Intelligence Research, 79, pp.1-40.
19. Santy, S., Liang, J.T., Bras, R.L., Reinecke, K. and Sap, M., 2023. NLPositionality:
Characterizing design biases of datasets and models. arXiv preprint arXiv:2306.01943.
20. Ernst, J.S., Marton, S., Brinkmann, J., Vellasques, E., Foucard, D., Kraemer, M. and Lambert,
M., 2023. Bias mitigation for large language models using adversarial learning. In CEUR Workshop
Proceedings (Vol. 3523, pp. 1-14). RWTH Aachen.
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
4.
HIGH LEVEL DESIGN
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
5. DATA PRE-PROCESSING
The figure shows the data preprocessing pipeline employed to transform the raw [Link] file
into a structured [Link] format for simpler analysis. The [Link] file has examples
grouped into two types: intersentence and intrasentence data. These are biases that cut across
multiple sentences or take place within one sentence, respectively. Each group has several data
entries with several keys, both necessary and supplementary information. In preprocessing, the data
is divided into these two categories, and from both of them, only the useful keys like context, target,
bias_type, gold_label, stereotype, anti-stereotype, and unrelated are kept. These fields are crucial for
model evaluation, bias analysis, and downstream tasks. The rest of the other keys, generally
metadata, are ignored. The cleaned and filtered data from both inter- and intra-sentence categories are
then combined and written into a CSV file called [Link]. This organized format enables easier
manipulation and analysis with the help of libraries such as Pandas and allows for effective input
preparation for machine learning models used for bias detection or mitigation.
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
6. IMPLEMENTATION PART - 1
This is a representation of the workflow to debias a language model based on the StereoSet dataset.
1. StereoSet Bias Categories
There are three evaluation categories offered by StereoSet:
A. Stereotype: Typical prejudiced or biased completion.
B. Anti-stereotype: A positive or neutral alternative.
C. Unrelated: Irrelevant or off-topic completion.
2. Model Evaluation
These categories are passed into a language model for the purpose of analyzing how the model
completes the prompt.
The result is indicative of the bias inclination of the model.
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
3. Biased Model Identification
Triggers are provided to a recognized biased or racist model.
Outputs are referred to as biased completions.
4. Debiasing Process
The biased completions go through the initial model (presumably a retrained or neutralized variant).
This operation should produce debiased completions that:
Erase hurtful stereotypes
Emphasize fairness and inclusiveness
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
This diagram represents the pipeline for detecting and assessing bias in language models using the
StereoSet dataset.
1. StereoSet Dataset
A test used to measure social biases (such as gender, race, profession) in language models.
Split into:
A. Intersentence bias: Bias across sentences.
B. Intrasentence bias: Bias in one sentence.
2. Train-Test Split
Intrasentence data is separated:
60% for training and 40% for testing
3. Test Set Processing
● The test set is divided evenly:
● 50% goes for testing test prompts
● 50% is utilized for sentence-level processing
4. Sentence Tokenization and Embedding
● Sentences are tokenized with NLTK.
● TF-IDF (Term Frequency-Inverse Document Frequency) is utilized to extract important
words like:
● black, man, bank
5. Word Association and Substitution
● For every word, a list of stereotypical or related words is produced.
● Sentences are altered by substituting original words with these related words.
● Several sentence variations are produced with substitutions.
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
IMPLEMENTATION PART - 2
Our bias-free Natural Language Generation (NLG) model was built using a structured, multi-layered
approach that prioritized fairness, contextual sensitivity, and high-quality text output.
We started by constructing a synthetic dataset composed of profession-based prompts (e.g., “The
doctor said that,” “The nurse decided to”) designed to simulate real-world bias scenarios. Each
sample was labeled to indicate whether it contained bias, allowing us to train and test the model in a
controlled environment.
For bias detection, we developed a custom module based on BERT. This module was trained to
classify text as biased or unbiased—not just based on obvious gender cues, but also by interpreting
subtle contextual clues within the embedding space. This allowed us to catch implicit biases that
traditional rule-based approaches often miss.
To actively reduce bias during text generation, we introduced an adversarial debiasing mechanism.
A discriminator network was trained to identify bias from the hidden layers of the T5 model. When
the discriminator flagged bias in generated text, a penalty was sent back to the T5 model,
encouraging it to revise its language patterns in real time. This adversarial feedback loop helped steer
the model toward more fair and inclusive generations.
We further refined this with Reinforcement Learning (RL). The T5 model was fine-tuned using a
custom reward function that considered both fluency and fairness. We combined the traditional
language modeling loss with feedback from the bias discriminator, adjusting the weight of the
fairness penalty carefully to avoid producing stiff or unnatural responses.
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
Training ran across multiple epochs to ensure the model had enough exposure to generalize while
still learning to avoid bias. We then evaluated the system using a predefined set of profession-based
prompts, comparing outputs from the original (biased) T5 model and our debiased version. The
results showed a clear reduction in biased language—especially male-skewed terms—though some
responses required minor fine-tuning for fluency.
To support scalability and reuse, we implemented model saving/loading features so the trained
system could be easily deployed or integrated into applications. Final evaluation included both
qualitative (manual review) and quantitative(bias frequency counting) assessments, confirming that
the model had learned to produce fairer and more balanced text.
In essence, our approach goes well beyond simple data balancing. It combines context-aware bias
detection, adversarial feedback, and reinforcement learning into a cohesive pipeline that trains
NLG systems to be more ethical, inclusive, and responsible—without sacrificing the quality of the
generated text.
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
7.
CONCLUSION OF CAPSTONE PROJECT PHASE – II
In this project, we are working to reduce societal biases in Natural Language Generation through the
use of adversarial debiasing and reinforcement learning. We have found and incorporated two
prominent benchmark datasets—StereoSet and BOLD—to aid in bias detection and testing.
Processing of the StereoSet dataset has been accomplished successfully, and processing of the BOLD
dataset is underway.
An initial iteration of our suggested model has already been put into practice, proving to be a
working and promising basis for continued development. In the future, we will further develop the
model, add fairness metrics, and test its efficacy on a wide range of text generation tasks to provide
inclusive and ethical AI systems.
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025
Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______
8.
PLAN OF WORK FOR CAPSTONE PROJECT PHASE - III
In Phase III of the capstone project, our main goals are:
● Finish at least 50% of model deployment, emphasizing bringing together adversarial
debiasing and reinforcement learning modules.
● Process and analyze the BOLD dataset, merging it with available StereoSet evaluations for
thorough bias detection.
● Release a review paper consolidating our survey of the literature, current bias mitigation
methods in generative AI, and initial findings from our implementation.
● Start developing test metrics to assess fairness, coherence, and contextual accuracy.
● Gear up for final integration, testing, and deployment in Phase IV.
_______________________________________________________________________________
______
Dept. of CSE(AI&ML) Jan-May 2025