0% found this document useful (0 votes)
17 views20 pages

Equitable Text Generation: Bias Mitigation

The dissertation titled 'Towards Equitable Text Generation: Bias Detection and Mitigation' presents a framework called CEDAR aimed at detecting and mitigating biases in Natural Language Generation (NLG) models. The project employs a multi-stage approach that includes a BERT-based bias detection module, adversarial training, and reinforcement learning to ensure the generation of fair and coherent text. The work addresses significant ethical concerns related to AI-generated content while striving for linguistic performance and responsibility.

Uploaded by

rahulr8310
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
17 views20 pages

Equitable Text Generation: Bias Mitigation

The dissertation titled 'Towards Equitable Text Generation: Bias Detection and Mitigation' presents a framework called CEDAR aimed at detecting and mitigating biases in Natural Language Generation (NLG) models. The project employs a multi-stage approach that includes a BERT-based bias detection module, adversarial training, and reinforcement learning to ensure the generation of fair and coherent text. The work addresses significant ethical concerns related to AI-generated content while striving for linguistic performance and responsibility.

Uploaded by

rahulr8310
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Dissertation on

“Towards Equitable Text Generation:


Bias Detection and Mitigation”

Submitted in partial fulfilment of the requirements for the award of the degree
of
Bachelor of Technology

In

CSE(AI&ML)

UE22AM320A – Capstone Project Phase - II

Submitted by:

Name - Kavya V SRN1- PES1UG22AM083


Name - Samprith Jagtap D SRN2- PES1UG22AM145
Name- Suraj B M SRN3- PES1UG22AM172
Name- T Jaiwanth SRN4- PES1UG22AM173

Under the guidance of

Prof. Bhaskar Jyoti Das


Adjunct Professor
PES University

January - May 2025

DEPARTMENT OF CSE(AI&ML)
FACULTY OF ENGINEERING
PES UNIVERSITY
(Established under Karnataka Act No. 16 of 2013)
100 feet Ring road, BSK 3rd stage, Hosakerehalli, Bengaluru – 560085
PES UNIVERSITY
(Established under Karnataka Act No. 16 of 2013)
100ft Ring Road, Bengaluru – 560 085, Karnataka, India

FACULTY OF ENGINEERING

CERTIFICATE
This is to certify that the dissertation entitled

‘Towards Equitable Text Generation:


Bias Detection and Mitigation’

is a bonafide work carried out by

Name- Kavya V PES1UG22AM083


Name- Samprith Jagtap D PES1UG22AM145
Name- Suraj B M PES1UG22AM172
Name- T Jaiwanth PES1UG22AM173

In partial fulfilment for the completion of Sixth-semester Capstone Project Phase - II


(UE22AM320A) in the Program of Study -Bachelor of Technology in CSE(AI&ML) under
rules and regulations of PES University, Bengaluru during the period January 2025 – May
2025. It is certified that all corrections/suggestions indicated for internal assessment have
been incorporated in the report. The dissertation has been approved as it satisfies the
6th-semester academic requirements in respect of project work.

Signature Signature Signature


Prof. Bhaskar Jyoti Das Dr. Jayashree R Dean, Engineering and
Adjunct Professor Chairperson Technology

External Viva

Name of the Examiners Signature with Date

1. __________________________ __________________________

2. __________________________ __________________________
DECLARATION

We hereby declare that the Capstone Project Phase - II entitled “Towards Equitable
Text Generation: Bias Detection and Mitigation” has been carried out by us under
the guidance of Prof. Bhaskar Jyoti Das, Professor and submitted in partial
fulfilment of the course requirements for the award of the degree of Bachelor of
Technology in CSE(AI&ML) of PES University, Bengaluru during the academic
semester January – May 2025. The matter embodied in this report has not been
submitted to any other university or institution for the award of any degree.

SRN1- PES1UG22AM083 Name - Kavya V


SRN2- PES1UG22AM145 Name - Samprith Jagtap D
SRN3- PES1UG22AM172 Name- Suraj B M
SRN4- PES1UG22AM173 Name- T Jaiwanth
ACKNOWLEDGEMENT

I would like to express my gratitude to Prof. Bhaskar Jyoti Das, Department of CSE(AI&ML),
PES University, for his/ her continuous guidance, assistance, and encouragement throughout the
development of UE22CS320A - Capstone Project Phase – II.

I am grateful to Capstone Project Coordinator, Prof. Shwetha K N for organizing, managing, and
helping with the entire process.

I take this opportunity to thank Dr. Jayashree R, Professor & Chairperson, Department of
CSE(AI&ML), PES University, for all the knowledge and support I have received from the
department. I would like to thank Dean, Engineering and Technology, PES University for his help.

I am deeply grateful to Prof. Jawahar Doreswamy, Chancellor, PES University, Dr. Suryaprasad J,
Vice-Chancellor, PES University, for providing me with various opportunities and enlightenment
every step of the way. Finally, Phase - II of the project could not have been completed without the
continual support and encouragement I have received from my family and friends.
ABSTRACT

Natural Language Generation (NLG) models have made impressive strides in producing text that
mimics human language. However, these models often inherit and amplify societal
biases—especially those related to gender, profession, and race—that are embedded in their training
data. As a result, AI-generated content can unintentionally reinforce harmful stereotypes, raising
serious ethical and social concerns. To address this, our project introduces CEDAR (Context-Aware
Embedding-Based Dynamic Adversarial Reinforcement learning), a novel, multi-stage framework
designed to detect and mitigate biases in NLG systems. CEDAR combines three key components: a
BERT-based bias detection module that identifies contextual bias, adversarial training that
dynamically reduces bias during model learning, and reinforcement learning with fairness-aware
reward functions to guide the generation of fluent and unbiased text. Our framework is designed to
strike a balance between ethical responsibility and linguistic performance—ensuring that generated
outputs remain coherent and semantically relevant while minimizing both explicit and implicit
biases. This work aims to advance the development of more responsible, inclusive, and trustworthy
AI systems.
​ ​ ​ ​ ​ ​ TABLE OF CONTENTS

Chapter No. Title Page No.

1.​ INTRODUCTION 01

2.​ PROBLEM DEFINITION 02

3.​ EXTENDED LITERATURE SURVEY 03


3.1 “Towards Robust NLG Bias Evaluation with Syntactically-diverse
Prompts”
3.2 “Societal Biases in Language Generation: Progress and Challenges”
3.3 “The Woman Worked as a Babysitter: On Biases in Language
Generation”
3.4 “This is a Problem, Don’t You Agree? Framing and Bias in Human
Evaluation for Natural Language Generation”
3.5 “Adhering, Steering, and Queering: Treatment of Gender in Natural
Language Generation”
3.6 “ Bias mitigation for large language models using adversarial
learning”
3.7 “NLPositionality: Characterizing design biases of datasets and
models”
3.8 “Undesirable biases in NLP: Addressing challenges of
measurement”
3.9 “Data augmentation techniques in natural language processing”

4.​ HIGH LEVEL DESIGN 05

5.​ DATA PRE-PROCESSING 06

6.​ IMPLEMENTATION 07

7.​ CONCLUSION OF CAPSTONE PROJECT PHASE – II 12

8.​ PLAN OF WORK FOR CAPSTONE PROJECT PHASE - III 13


​ ​​ ​ ​ ​ LIST OF FIGURES

Figure No. Title Page


No.

Figure- 1 High Level Design 5

Figure -2 ​ Data Pre-Processing ​ ​ ​ ​ 6

Figure - 3​ ​ Implementation (5.a)​ ​ ​ ​ 7

Figure - 4​ ​ Implementation (5.b)​ ​ ​ ​ 8


CHAPTER 1

INTRODUCTION

1.1
We implemented our bias-free NLG model using a layered approach focused on fairness, context,
and text quality. First, we built a synthetic dataset of profession-based prompts labeled for bias
presence, simulating real-world scenarios. A BERT-based classifier was trained to detect both
explicit and subtle biases by understanding contextual embeddings. Then, we introduced adversarial
debiasing using a discriminator network that flagged biased outputs from the T5 decoder. When bias
was detected, a penalty was backpropagated to guide the model toward fairer language. To further
improve fairness, we fine-tuned the T5 model using reinforcement learning with a custom reward
function combining fluency and fairness. By carefully tuning the penalty weights, we preserved text
coherence while reducing biased outputs. The model was trained over multiple epochs and evaluated
using a fixed set of prompts. We observed significant bias reduction compared to the original model,
with minimal impact on output fluency. The entire pipeline, including model saving and reloading,
ensures repeatability and integration readiness. Overall, our system combines context-aware
detection, adversarial feedback, and reinforcement learning to produce fair, fluent, and
responsible text.

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025 1


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

2.
PROBLEM DEFINITION

NLG models create biased content because of unbalanced data and model limitations. The biases
perpetuate harmful stereotypes and lower the credibility of AI.

The task is to:


Identify and minimize implicit and explicit biases in NLG models without compromising text quality
and contextual appropriateness.
This demands to tackle:
●​ Subtle bias detection,
●​ Fairness without overcorrection,
●​ High computational cost,
●​ Subjectivity in defining fairness.

Our method introduces a strong, scalable framework to provide fair and ethical AI text generation.

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

3. EXTENDED LITERATURE SURVEY


1. Aggarwal, A., Sun, J. and Peng, N., 2022. Towards robust NLG bias evaluation with
syntactically-diverse prompts. arXiv preprint arXiv:2212.01700.

2. Sheng, E., Chang, K.W., Natarajan, P. and Peng, N., 2021. Societal biases in language
generation: Progress and challenges. arXiv preprint arXiv:2105.04054.

3. Sheng, E., Chang, K.W., Natarajan, P. and Peng, N., 2019. The woman worked as a babysitter:
On biases in language generation. arXiv preprint arXiv:1909.01326.

4. ​ Schoch, S., Yang, D. and Ji, Y., 2020, December. “This is a Problem, Don’t You Agree?”
Framing and Bias in Human Evaluation for Natural Language Generation. In Proceedings of the 1st
Workshop on Evaluating NLG Evaluation (pp. 10-16).

5. Strengers, Y., Qu, L., Xu, Q. and Knibbe, J., 2020, April. Adhering, steering, and queering:
Treatment of gender in natural language generation. In Proceedings of the 2020 CHI Conference on
Human Factors in Computing Systems (pp. 1-14).

6. ​ Dhamala, J., Sun, T., Kumar, V., Krishna, S., Pruksachatkun, Y., Chang, K.W. and
Gupta, R., 2021, March. Bold: Dataset and metrics for measuring biases in open-ended language
generation. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency
(pp. 862-872).

7. ​ Wei, J.T.Z. and Jia, R., 2021. The statistical advantage of automatic NLG metrics at the
system level. arXiv preprint arXiv:2105.12437.

8. ​ Das, M. and Balke, W.T., 2022, October. Quantifying bias from decoding techniques in
natural language generation. In Proceedings of the 29th International Conference on Computational
Linguistics (pp. 1311-1323).

9. ​ Zhou, K., Blodgett, S.L., Trischler, A., Daumé III, H., Suleman, K. and Olteanu, A., 2022.
Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications. arXiv
preprint arXiv:2205.06828.

10. Sahoo, N., Gupta, H. and Bhattacharyya, P., 2022. Detecting unintended social bias in toxic
language datasets. arXiv preprint arXiv:2210.11762.

11. ​ Dhamala, J., Sun, T., Kumar, V., Krishna, S., Pruksachatkun, Y., Chang, K.W. and Gupta, R.,
_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

2021, March. Bold: Dataset and metrics for measuring biases in open-ended language generation. In
Proceedings of the 2021 ACM conference on fairness, accountability, and transparency (pp.
862-872).

12. ​ Wei, J.T.Z. and Jia, R., 2021. The statistical advantage of automatic NLG metrics at the
system level. arXiv preprint arXiv:2105.12437.

13. ​ Das, M. and Balke, W.T., 2022, October. Quantifying bias from decoding techniques in
natural language generation. In Proceedings of the 29th International Conference on Computational
Linguistics (pp. 1311-1323).

14. ​ Zhou, K., Blodgett, S.L., Trischler, A., Daumé III, H., Suleman, K. and Olteanu, A., 2022.
Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications. arXiv
preprint arXiv:2205.06828.

15. Sahoo, N., Gupta, H. and Bhattacharyya, P., 2022. Detecting unintended social bias in toxic
language datasets. arXiv preprint arXiv:2210.11762.

16 ​ Guo, Y., Guo, M., Su, J., Yang, Z., Zhu, M., Li, H., Qiu, M. and Liu, S.S., 2024. Bias in large
language models: Origin, evaluation, and mitigation. arXiv preprint arXiv:2411.10915.

17. ​ Pellicer, L.F.A.O., Ferreira, T.M. and Costa, A.H.R., 2023. Data augmentation techniques in
natural language processing. Applied Soft Computing, 132, p.109803.

18. ​ Van der Wal, O., Bachmann, D., Leidinger, A., van Maanen, L., Zuidema, W. and Schulz, K.,
2024. Undesirable biases in NLP: Addressing challenges of measurement. Journal of Artificial
Intelligence Research, 79, pp.1-40.

19. ​ Santy, S., Liang, J.T., Bras, R.L., Reinecke, K. and Sap, M., 2023. NLPositionality:
Characterizing design biases of datasets and models. arXiv preprint arXiv:2306.01943.

20. ​ Ernst, J.S., Marton, S., Brinkmann, J., Vellasques, E., Foucard, D., Kraemer, M. and Lambert,
M., 2023. Bias mitigation for large language models using adversarial learning. In CEUR Workshop
Proceedings (Vol. 3523, pp. 1-14). RWTH Aachen.

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

4.
HIGH LEVEL DESIGN

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

5. DATA PRE-PROCESSING

The figure shows the data preprocessing pipeline employed to transform the raw [Link] file
into a structured [Link] format for simpler analysis. The [Link] file has examples
grouped into two types: intersentence and intrasentence data. These are biases that cut across
multiple sentences or take place within one sentence, respectively. Each group has several data
entries with several keys, both necessary and supplementary information. In preprocessing, the data
is divided into these two categories, and from both of them, only the useful keys like context, target,
bias_type, gold_label, stereotype, anti-stereotype, and unrelated are kept. These fields are crucial for
model evaluation, bias analysis, and downstream tasks. The rest of the other keys, generally
metadata, are ignored. The cleaned and filtered data from both inter- and intra-sentence categories are
then combined and written into a CSV file called [Link]. This organized format enables easier
manipulation and analysis with the help of libraries such as Pandas and allows for effective input
preparation for machine learning models used for bias detection or mitigation.

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

6. IMPLEMENTATION PART - 1

This is a representation of the workflow to debias a language model based on the StereoSet dataset.
1. StereoSet Bias Categories
There are three evaluation categories offered by StereoSet:
A.​ Stereotype: Typical prejudiced or biased completion.
B.​ Anti-stereotype: A positive or neutral alternative.
C.​ Unrelated: Irrelevant or off-topic completion.
2. Model Evaluation
These categories are passed into a language model for the purpose of analyzing how the model
completes the prompt.
The result is indicative of the bias inclination of the model.

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

3. Biased Model Identification


Triggers are provided to a recognized biased or racist model.
Outputs are referred to as biased completions.
4. Debiasing Process
The biased completions go through the initial model (presumably a retrained or neutralized variant).
This operation should produce debiased completions that:
Erase hurtful stereotypes
Emphasize fairness and inclusiveness

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

This diagram represents the pipeline for detecting and assessing bias in language models using the
StereoSet dataset.
1. StereoSet Dataset
A test used to measure social biases (such as gender, race, profession) in language models.
Split into:
A.​ Intersentence bias: Bias across sentences.
B.​ Intrasentence bias: Bias in one sentence.
2. Train-Test Split
Intrasentence data is separated:
60% for training and 40% for testing
3. Test Set Processing
●​ The test set is divided evenly:
●​ 50% goes for testing test prompts
●​ 50% is utilized for sentence-level processing
4. Sentence Tokenization and Embedding
●​ Sentences are tokenized with NLTK.
●​ TF-IDF (Term Frequency-Inverse Document Frequency) is utilized to extract important
words like:
●​ black, man, bank
5. Word Association and Substitution
●​ For every word, a list of stereotypical or related words is produced.
●​ Sentences are altered by substituting original words with these related words.
●​ Several sentence variations are produced with substitutions.

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

IMPLEMENTATION PART - 2

Our bias-free Natural Language Generation (NLG) model was built using a structured, multi-layered
approach that prioritized fairness, contextual sensitivity, and high-quality text output.

We started by constructing a synthetic dataset composed of profession-based prompts (e.g., “The


doctor said that,” “The nurse decided to”) designed to simulate real-world bias scenarios. Each
sample was labeled to indicate whether it contained bias, allowing us to train and test the model in a
controlled environment.

For bias detection, we developed a custom module based on BERT. This module was trained to
classify text as biased or unbiased—not just based on obvious gender cues, but also by interpreting
subtle contextual clues within the embedding space. This allowed us to catch implicit biases that
traditional rule-based approaches often miss.

To actively reduce bias during text generation, we introduced an adversarial debiasing mechanism.
A discriminator network was trained to identify bias from the hidden layers of the T5 model. When
the discriminator flagged bias in generated text, a penalty was sent back to the T5 model,
encouraging it to revise its language patterns in real time. This adversarial feedback loop helped steer
the model toward more fair and inclusive generations.

We further refined this with Reinforcement Learning (RL). The T5 model was fine-tuned using a
custom reward function that considered both fluency and fairness. We combined the traditional
language modeling loss with feedback from the bias discriminator, adjusting the weight of the
fairness penalty carefully to avoid producing stiff or unnatural responses.

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

Training ran across multiple epochs to ensure the model had enough exposure to generalize while
still learning to avoid bias. We then evaluated the system using a predefined set of profession-based
prompts, comparing outputs from the original (biased) T5 model and our debiased version. The
results showed a clear reduction in biased language—especially male-skewed terms—though some
responses required minor fine-tuning for fluency.

To support scalability and reuse, we implemented model saving/loading features so the trained
system could be easily deployed or integrated into applications. Final evaluation included both
qualitative (manual review) and quantitative(bias frequency counting) assessments, confirming that
the model had learned to produce fairer and more balanced text.

In essence, our approach goes well beyond simple data balancing. It combines context-aware bias
detection, adversarial feedback, and reinforcement learning into a cohesive pipeline that trains
NLG systems to be more ethical, inclusive, and responsible—without sacrificing the quality of the
generated text.

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

7.
CONCLUSION OF CAPSTONE PROJECT PHASE – II

In this project, we are working to reduce societal biases in Natural Language Generation through the
use of adversarial debiasing and reinforcement learning. We have found and incorporated two
prominent benchmark datasets—StereoSet and BOLD—to aid in bias detection and testing.
Processing of the StereoSet dataset has been accomplished successfully, and processing of the BOLD
dataset is underway.

An initial iteration of our suggested model has already been put into practice, proving to be a
working and promising basis for continued development. In the future, we will further develop the
model, add fairness metrics, and test its efficacy on a wide range of text generation tasks to provide
inclusive and ethical AI systems.

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025


Towards Equitable Text Generation:Bias Detection and Mitigation
_______________________________________________________________________________
______

8.
PLAN OF WORK FOR CAPSTONE PROJECT PHASE - III

In Phase III of the capstone project, our main goals are:


●​ Finish at least 50% of model deployment, emphasizing bringing together adversarial
debiasing and reinforcement learning modules.
●​ Process and analyze the BOLD dataset, merging it with available StereoSet evaluations for
thorough bias detection.
●​ Release a review paper consolidating our survey of the literature, current bias mitigation
methods in generative AI, and initial findings from our implementation.
●​ Start developing test metrics to assess fairness, coherence, and contextual accuracy.
●​ Gear up for final integration, testing, and deployment in Phase IV.

_______________________________________________________________________________
______

Dept. of CSE(AI&ML) Jan-May 2025

Common questions

Powered by AI

Reinforcement learning is used by fine-tuning the T5 model with a custom reward function combining fluency and fairness. This approach integrates traditional language modeling loss with feedback from the bias discriminator, adjusting the fairness penalty weight to maintain text coherence while reducing bias .

One specific way the project modifies sentence data is through word association and substitution, wherein sentences are altered by substituting original words with stereotypical or related words to generate several variations and assess model bias .

The fairness and fluency of the debiased NLG model are assessed qualitatively through manual reviews and quantitatively by counting bias frequency in generated outputs, utilizing predefined profession-based prompts for comparison .

The project addresses challenges such as detecting subtle biases, ensuring fairness without overcorrection, managing the high computational cost, and dealing with the subjectivity inherent in defining fairness .

The project introduces CEDAR (Context-Aware Embedding-Based Dynamic Adversarial Reinforcement learning), a multi-stage framework for bias detection and mitigation in NLG systems. It combines a BERT-based bias detection module, adversarial training, and reinforcement learning with fairness-aware reward functions .

The BERT-based bias detection module identifies contextual biases by understanding contextual embeddings. It detects both explicit and subtle biases, allowing the system to catch implicit biases that traditional methods may overlook .

The StereoSet dataset evaluates biases in three categories: stereotype (typical prejudiced completion), anti-stereotype (positive or neutral alternative), and unrelated (irrelevant or off-topic completion).

The project ensures repeatability and integration readiness by implementing model saving/loading features, allowing the trained system to be easily deployed or integrated into applications .

Adversarial training involves a discriminator network that flags biased outputs from the T5 decoder. When bias is detected, a penalty is backpropagated to guide the model towards using fairer language, thus reducing bias during text generation .

The project balances ethical responsibility and linguistic performance by ensuring the generation of fluent and semantically relevant text while minimizing explicit and implicit biases through a combination of bias detection, adversarial feedback, and reinforcement learning .

You might also like