0% found this document useful (0 votes)
24 views24 pages

Indicsynth: A Large-Scale Multilingual Synthetic Speech Dataset For Low-Resource Indian Languages

IndicSynth is a large-scale multilingual synthetic speech dataset containing approximately 4,000 hours of synthetic speech from 989 speakers across 12 low-resourced Indian languages. The dataset aims to address the linguistic bias in audio deepfake detection and anti-spoofing models, which often perform poorly on out-of-domain languages. IndicSynth is partitioned into mimicry and diversity subsets to enhance its utility for research, and it is accessible for further exploration.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
24 views24 pages

Indicsynth: A Large-Scale Multilingual Synthetic Speech Dataset For Low-Resource Indian Languages

IndicSynth is a large-scale multilingual synthetic speech dataset containing approximately 4,000 hours of synthetic speech from 989 speakers across 12 low-resourced Indian languages. The dataset aims to address the linguistic bias in audio deepfake detection and anti-spoofing models, which often perform poorly on out-of-domain languages. IndicSynth is partitioned into mimicry and diversity subsets to enhance its utility for research, and it is accessible for further exploration.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

IndicSynth: A Large-Scale Multilingual Synthetic Speech Dataset for

Low-Resource Indian Languages

Divya V. Sharma Vijval Ekbote Anubha Gupta


SBILab, IIIT-Delhi SBILab, IIIT-Delhi SBILab, IIIT-Delhi
divyas@[Link] vijval22569@[Link] anubha@[Link]

Abstract et al., 2024; Yi et al., 2023). Despite these ap-


plications, there exists a severe threat of misuse
Recent advances in synthetic speech generation of these technologies to spread fake news, malign
technology have facilitated the generation of personalities, financial fraud, or other criminal ac-
high-quality synthetic (fake) speech that emu- tivities (Müller et al., 2024b; Ju et al., 2024). It is
lates human voices. These technologies pose
because these technologies can generate realistic
a threat of misuse for identity theft and the
spread of misinformation. Consequently, the synthetic speech recordings that can deceive both
misuse of such powerful technologies necessi- humans and security systems, such as speaker veri-
tates the development of robust and generaliz- fication. Therefore, to prevent potential misuse, it
able audio deepfake detection (ADD) and anti- is crucial to develop robust systems to detect fake
spoofing models. However, such models are (synthetic) audio generated by these technologies
often linguistically biased. Consequently, the (Müller et al., 2024b).
models trained on datasets in one language ex- An audio deepfake is a synthetic speech record-
hibit a low accuracy when evaluated on out-of-
ing that appears natural enough to deceive humans
domain languages. Such biases reduce the us-
ability of these models and highlight the urgent and can potentially be used to spread misinforma-
need for multilingual synthetic speech datasets tion that can cause public unrest and panic (Müller
for bias mitigation research. However, most et al., 2024b). In contrast, audio spoofing refers
available datasets are in English or Chinese. to the techniques used to generate synthetic au-
The dearth of multilingual synthetic datasets dios that mimic a target speaker’s voice. Such
hinders multilingual ADD and anti-spoofing techniques are often misused to deceive biomet-
research. Furthermore, the problem intensi-
ric security systems, such as speaker verification
fies in countries with rich linguistic diversity,
such as India. Therefore, we introduce Indic- (Kawa et al., 2022; Müller et al., 2024b). Conse-
Synth, which contains 4,000 hours of synthetic quently, audio spoofing can lead to impersonation
speech from 989 target speakers, including 456 and identity theft. For instance, in 2019, imposters
females and 533 males for 12 low-resourced In- used spoofed audios of a corporate executive for fi-
dian languages. The dataset includes rich meta- nancial fraud of 243,000 US dollars (USD) (Munir
data covering gender details and target speaker et al., 2024; Frank and Schönherr, 2021). Simi-
identifiers. Experimental results demonstrate larly, an audio deepfake-based cybercrime caused a
that IndicSynth is a valuable contribution to
loss of 35 million USD for a UAE-based company
multilingual ADD and anti-spoofing research.
The dataset can be accessed from https:// (Rabhi et al., 2024). Audio deepfakes are often
[Link]/vdivyas/IndicSynth. misused to tarnish the reputation of celebrities and
politicians and to manipulate public opinion during
1 Introduction elections (Kawa et al., 2022). Therefore, devel-
oping robust audio deepfake detection and anti-
Recent advances in text-to-speech (TTS) and voice spoofing technologies is crucial for social good.
conversion (VC) models have enabled the gener- For the development of robust audio deepfake
ation of high-quality synthetic speech recordings detection (ADD) models, realistic synthetic speech
that emulate human voices (Ba et al., 2023). Such datasets are vital. However, most existing datasets
recordings have applications in diverse domains, are in high-resource languages, such as English
including assistive technologies, media, entertain- and Chinese (Müller et al., 2024b; Ba et al., 2023;
ment, and education (Bird and Lotfi, 2023; Rabhi Munir et al., 2024; Frank and Schönherr, 2021;
22037
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 22037–22060
July 27 - August 1, 2025 ©2025 Association for Computational Linguistics
Khalid et al., 2021). Consequently, due to the ease 2. We evaluate the linguistic bias in state-of-the-
of availability of datasets, most previous works on art audio deepfake detection (ADD) models
ADD focus on these languages (Yi et al., 2024). and the vulnerability of SOTA speaker verifi-
However, the ADD models trained on synthetic cation models to impersonation attacks from
speech datasets in one language often exhibit a sig- multilingual audio spoofs. Consequently, we
nificantly lower accuracy when evaluated on out- investigate IndicSynth’s utility towards gener-
of-domain languages (Müller et al., 2024b). There- alizable ADD and anti-spoofing.
fore, synthetic speech datasets in low-resource lan-
guages are urgently needed to enhance the global 3. We qualitatively and quantitatively assess the
usability of these models. The need for such linguistic authenticity of our dataset through
datasets intensifies in countries with rich linguistic t-SNE plots and a state-of-the-art language
diversity, such as India. India has 22 constitution- identification model.
ally recognized languages spoken by one billion
2 Related Works
speakers, and around 75% of Indians come across
some deepfake content in a year (Javed et al., 2023; Audio DeepFake Detection and Anti-Spoofing:
T and T, 2024). The proliferation of audio deepfakes and audio
In this work, we introduce IndicSynth, a novel spoofs-related fraud prompted research communi-
large-scale multilingual synthetic speech dataset ties to organize Audio DeepFake Detection (ADD)
for 12 low-resourced Indian languages. IndicSynth and ASVspoof challenges (Yi et al., 2022, 2023;
contains approximately 4,000 hours of synthetic Wang et al., 2020; Liu et al., 2023). Despite these
speech from 989 target speakers, including 456 fe- initiatives, a critical yet underexplored challenge
males and 533 males. The dataset includes rich is the lack of generalizability of ADD and anti-
metadata covering the gender and identifiers of spoofing models to out-of-domain scenarios (Kor-
the target speakers. Additionally, to enhance In- shunov and Marcel, 2022; Yousif et al., 2024; Xie
dicSynth’s utility for multilingual audio deepfake et al., 2024; Kawa et al., 2022; Müller et al., 2024a).
detection (ADD) and anti-spoofing research, we ADD and anti-spoofing models trained on speech
partition our dataset into mimicry and diversity datasets in one language exhibit significantly re-
subsets. The mimicry subset contains synthetic au- duced accuracy when evaluated on out-of-domain
dios closely mimicking the bonafide target voices. languages. Such linguistic biases reduce the util-
In contrast, the diversity subset contains a more di- ity of these models (Sharma and Buduru, 2022;
verse set of realistic synthetic voices. To generate Sharma, 2024).
IndicSynth, we apply state-of-the-art (SOTA) TTS Dearth of Datasets: To mitigate linguistic bi-
and VC models on a publicly available bonafide ases in ADD and anti-spoofing, researchers have
(real) speech dataset. After generation, we inves- explored domain adaptation and data augmenta-
tigate whether SOTA ADD models can accurately tion techniques (Ba et al., 2023; Xie et al., 2024).
classify IndicSynth’s synthetic audio as fake. Next, However, these techniques require synthetic speech
we evaluate the linguistic authenticity of our dataset datasets in the target languages. Most publicly
using a SOTA language identification model. Sub- available synthetic datasets, such as FakeAVCeleb,
sequently, we evaluate whether the IndicSynth’s the ASVspoof 2019 dataset, and the ADD 2023
mimicry subset can deceive SOTA speaker verifica- challenge dataset, are in English or Chinese (Yi
tion models through impersonation attacks. For the et al., 2023; Müller et al., 2024b; Munir et al., 2024;
ease of reproducibility of our results, we use only Ba et al., 2023; Khalid et al., 2021; Wang et al.,
publicly available models for experimentation. 2020). Consequently, researchers introduced the
Key contributions of this study are: Urdu audio deepfake detection dataset that contains
16,830 spoofed audios (Munir et al., 2024). Simi-
1. We introduce IndicSynth, a novel multilingual larly, the WaveFake dataset containing 196 hours of
synthetic speech dataset for 12 low-resourced synthetic audios in English and Japanese was intro-
Indian languages containing approximately duced (Frank and Schönherr, 2021). Additionally,
4,000 hours of audio from 989 target speakers, the MLAAD dataset and the MLADDC datasets
including 456 females and 533 males. We were introduced (Müller et al., 2024b; SHAH et al.,
partition the dataset into mimicry and diversity 2024). However, these datasets do not include
subsets. gender details and target speaker identifiers. Gen-
22038
der details are essential for studies related to gen- target speech recording (v tgt ) and a transcript (ttxt )
der bias in ADD (Xu et al., 2024; Ju et al., 2024; as inputs. Subsequently, these models generate
Kawa et al., 2022; Haut et al., 2022). The target a synthetic speech recording (v tgt tts ) as output, as
speaker identifiers are crucial for developing de- illustrated Equation 1:
fense mechanisms against impersonation attacks.
In addition to these datasets, the MADD dataset v tgt tts = T T S(ttxt , v tgt ) (1)
contains 155.66 hours of synthetic audio for six
languages (Qi et al., 2024). Thus, to address the
dearth of publicly available large-scale multilin- IndicSuperb
gual synthetic speech datasets containing gender 1
Transcripts Bonafide Audios
information and identifiers of the target speakers,
ttxt
we introduce the IndicSynth. "भारत में हर भाषा
का सम्मान किया
"भारत में हर जाता है।"
भाषा का सम्मान Target (vtgt) Source (vsrc) Target (vtgt)
3 IndicSynth: Generation and Overview 2
किया जाता है।"

This study introduces IndicSynth, a novel large- Text-to-Speech Voice Conversion


scale multilingual synthetic speech dataset. In- "भारत में हर
भाषा का सम्मान
dicSynth contains approximately 4,000 hours of 3
Synthetic (vtgttts) किया जाता है।" Synthetic (vtgtvc)

speech recordings for 12 low-resourced target lan- IndicSynth


guages: Bengali, Gujarati, Hindi, Kannada, Malay-
alam, Marathi, Odia, Punjabi, Sanskrit, Tamil, Tel- Figure 2: IndicSynth’s generation methodology. We
ugu, and Urdu, as illustrated in Figure 1. This apply publicly available text-to-speech and voice con-
section describes IndicSynth’s generation method- version models to the publicly available bonafide Indic-
ology and its statistical details. Superb dataset to generate IndicSynth. IndicSuperb is
licensed under CC0 ("no rights reserved").
552.4
500 Females
Total Duration (hours)

Males 446.9 464.3 The generated synthetic speech (v tgt tts ) articu-
400
lates the transcript (ttxt ) while attempting to mimic
300 270.9
209.7 207.3 218.6 the target voice (v tgt ). In contrast to the TTS mod-
200 169.0 172.1 185.1
170.0
96.2 86.4110.5
170.1 els, voice conversion (VC) models take a bonafide
100 78.3 84.6 64.981.3
55.2 58.6 58.236.8
13.3 source speech recording (v src ) and a bonafide tar-
0
get speech recording (v tgt ) as inputs. Subsequently,
Gu ali
ati
Ka di
lay a
Ma m
hi
Pu ia
Sa abi
t
Tel l
u
u
i
kri
Tam
Ma nad

ug
Urd
Od
Hin

rat
ala
ng
jar

nj
ns

these models generate a synthetic speech recording


Be

Languages (v tgt vc ) as output, as illustrated in Equation 2:


Figure 1: Total duration (in hours) of synthetic male and
female voices in IndicSynth for each target language. v tgt vc = V C(v src , v tgt ) (2)
IndicSynth Generation Methodology: Firstly, The generated synthetic speech (v tgt vc ) articu-
we selected the IndicSuperb dataset to create the lates the transcript of the source audio (v src ), while
IndicSynth (Javed et al., 2023). IndicSuperb is attempting to mimic the target voice (v tgt ).
licensed under the Creative Commons CC0 li- Mimicry and Diversity: IndicSynth includes
cense ("no rights reserved") agreement. It con- two types of synthetic data, illustrated in Table 1:
tains bonafide (real) speech recordings and their
associated transcripts for the 12 target languages, 1. Mimicry: The mimicry subset contains syn-
as shown in Figure 2. Following the bonafide thetic audios that closely mimic target voices.
dataset selection, the next step is to apply publicly This subset is valuable for assessing the vul-
available text-to-speech (TTS) based voice cloning nerability of speaker verification systems to
models and voice conversion (VC) models to the impersonation attacks (Munir et al., 2024).
bonafide IndicSuperb data1 . TTS and VC mod- 2. Diversity: The diversity subset includes syn-
els are widely used for synthetic data generation thetic audios with low voice similarity to tar-
(Zhu et al., 2024). TTS models take a bonafide get voices. Consequently, this subset has more
1
diversity in synthetic voices. The diversity
Models used for IndicSynth generation: https://
[Link]/coqui-ai/TTS. The training datasets of these subset is valuable because training audio deep-
models are not fully disclosed. fake detection models on such diverse, mul-
22039
Language Model Category #Females #Male #Clips Duration (hrs)
XTTS-v2 Mimicry 18 10 28,056 50.67
Bengali FreeVC24 Diversity 18 10 27,336 51.46
VITS Diversity 18 10 28,056 49.30
XTTS-v2 Mimicry 25 34 59,118 97.22
Gujarati
FreeVC24 Diversity 25 34 59,660 99.70
XTTS-v2 Diversity 53 48 101,202 171.36
Hindi
FreeVC24 Diversity 53 48 104,736 167.77
XTTS-v2 Mimicry 13 43 55,611 127.44
Kannada
FreeVC24 Diversity 16 43 59,412 140.89
XTTS-v2 Mimicry 10 7 17,034 46.35
Malayalam
FreeVC24 Diversity 10 7 17,094 48.61
XTTS-v2 Mimicry 51 72 123,246 231.74
Marathi
FreeVC24 Diversity 51 72 130,150 246.461
XTTS-v2 Mimicry 22 4 26,052 45.34
Odia
FreeVC24 Diversity 22 4 26,184 46.30
XTTS-v2 Mimicry 67 55 122,244 191.60
Punjabi
FreeVC24 Diversity 67 55 126,110 199.11
XTTS-v2 Diversity 100 85 185,370 422.862
Sanskrit
FreeVC24 Diversity 100 85 192,134 576.21
XTTS-v2 Mimicry 32 106 138,276 280.42
Tamil
FreeVC24 Diversity 32 106 144,036 298.42
XTTS-v2 Mimicry 41 43 84,168 175.65
Telugu
FreeVC24 Diversity 41 43 85,728 179.41
XTTS-v2 Mimicry 21 26 47,094 72.34
Urdu
FreeVC24 Diversity 21 26 47,804 73.95
Table 1: Overview of IndicSynth, including generative model name, subset type (category), number of male and
female target speakers, number of audio clips, and total duration of synthetic audio (in hours) for each language.

tilingual synthetic datasets can enhance their were the same for high-quality voice conversion.
generalizability to out-of-domain languages. IndicSynth Metadata: IndicSynth contains sep-
arate metadata files for voice cloning done through
IndicSynth Generation: To generate synthetic each model for each target language3 . For the TTS
data for the mimicry subset, we fine-tuned the models, the metadata includes the target speaker
XTTS-v2 model on IndicSuperb for each of the fol- ID, the ID of the bonafide target voice sample, the
lowing 10 target languages: Bengali, Gujarati, Kan- target speaker’s gender, the transcript, and the ID
nada, Malayalam, Marathi, Odia, Punjabi, Tamil, of the generated synthetic audio clip. Similarly, for
Telugu, and Urdu2 . Fine-tuning XTTS-v2 and the VC models, the metadata includes the source
generating synthetic data using the same bonafide speaker ID, bonafide IndicSuperb source audio clip
dataset (IndicSuperb) helps ensure a high similarity ID, target speaker ID, bonafide IndicSuperb target
between synthetic and target voices. In contrast, audio clip ID, the gender of the speakers, and the ID
the diversity subset includes synthetic audios di- of the synthetic audio clip. Such metadata is also
rectly generated from TTS and VC models without valuable for studying gender bias in multilingual au-
fine-tuning on IndicSuperb. We utilized the pub- dio deepfake detection (ADD) (Xu et al., 2024; Ju
licly available Coqui VITS model (a TTS model et al., 2024). Overall, utilizing the bonafide Indic-
trained on undisclosed Bengali data) to generate Superb with the synthetic IndicSynth can facilitate
Bengali synthetic data for diversity subset (Eren multilingual ADD and anti-spoofing research.
and The Coqui TTS Team, 2021). Additionally, we
generated synthetic data for each of the 12 target 4 Evaluation of IndicSynth
languages using the publicly available XTTS-v2
4.1 IndicSynth for Audio DeepFake Detection
text-to-speech model and the FreeVC24 voice con-
version model (Eren and The Coqui TTS Team, Audio deepfake detection (ADD) models accept
2021), as illustrated in Table 1. The bonafide audios a speech recording as input and determine if it is
and transcripts were randomly chosen for Indic- bonafide (real) or synthetic (fake). These models
Synth creation. Also, we ensured that the genders are vital to prevent the spread of fake news or misin-
of the randomly chosen source and target speakers formation from synthetic audio. Thus, we urgently
need ADD models that are inclusive and general-
2
Code for fine-tuning XTTS-v2: [Link]
3
anhnh2002/XTTSv2-Finetuning-for-New-Languages IndicSynth’s directory structure is in Appendix (A).

22040
izable to unseen languages. A crucial first step in EER (%)
Language G. Model Aasist Aasist-L RawNet-2
developing generalizable ADD models is bench- XTTS-v2 70.125 56.150 56.737
marking state-of-the-art (SOTA) models on unseen Bengali FreeVC24 87.963 86.563 53.537
language datasets. Benchmarking can help inves- VITS 93.200 89.363 48.287
XTTS-v2 65.050 55.150 50.113
tigate potential biases in these models. Therefore, Gujarati
FreeVC24 86.163 86.888 53.425
we used IndicSynth to benchmark three state-of- Hindi
XTTS-v2 42.013 45.438 14.525
the-art (SOTA) publicly available ADD models: FreeVC24 81.775 81.913 48.513
XTTS-v2 55.563 49.188 42.425
Aasist, Aasist-L, and RawNet2 (Jung et al., 2022; Kannada
FreeVC24 73.30 78.950 50.874
Tak et al., 2021)4 . The benchmark ADD models are Malayalam
XTTS-v2 67.575 55.013 46.888
trained on an English dataset (LA partition of the FreeVC24 85.825 83.600 55.675
XTTS-v2 56.712 52.825 48.037
ASVspoof 2019 challenge) (Wang et al., 2020). We Marathi
FreeVC24 79.512 81.587 52.525
evaluated them on IndicSynth without fine-tuning. Odia
XTTS-v2 57.488 51.575 48.487
Setup: We created separate test sets for each FreeVC24 78.888 82.975 44.350
XTTS-v2 57.575 52.863 47.225
target language and generative model, as illustrated Punjabi
FreeVC24 81.925 82.225 53.775
in Table 2. Each set includes randomly chosen Sanskrit
XTTS-v2 33.438 38.775 7.95
4000 bonafide female voice samples, 4000 bonafide FreeVC24 84.238 86.563 58.35
XTTS-v2 61.725 51.188 51.650
male voice samples, 4000 synthetic female voice Tamil
FreeVC24 81.700 83.138 54.999
samples, and 4000 synthetic male voice samples. Telugu
XTTS-v2 54.275 52.700 46.275
The bonafide data was taken from IndicSuperb, FreeVC24 75.650 79.000 52.787
XTTS-v2 62.763 55.363 49.250
whereas synthetic data was taken from IndicSynth. Urdu
FreeVC24 78.088 79.825 50.438
We refer to them as IndicSynth-IndicSuperb sets. Table 2: Benchmarking audio deepfake detection (ADD)
Evaluation Metric: False Match Rate (FMR) models in IndicSynth-IndicSuperb test sets without do-
and False Non-Match Rate (FNMR) are widely main adaptation. For a given target language and a
used metrics for evaluating biometric systems. particular ADD model, the highest Equal Error Rate
FMR is the rate at which an ADD model incorrectly (EER%) achieved across various generative models is
highlighted in bold. When evaluated without domain
classifies synthetic audios as bonafide. In contrast,
adaptation, the benchmark ADD models achieve ele-
FNMR is the rate at which an ADD model incor- vated EER% on IndicSynth-IndicSuperb test sets. Train-
rectly classifies bonafide audios as synthetic. The ing ADD models on multilingual synthetic datasets,
FMR and FNMR values of an ADD model vary such as IndicSynth, can enhance their generalizability.
with classification thresholds. At a particular classi-
fication threshold, FMR becomes equal to FNMR.
The value of FMR when it becomes equal to the
FNMR is known as the Equal Error Rate (EER).
Equal Error Rate (EER) is a standard evaluation
metric for audio deepfake detection systems (Wang
et al., 2020; Liu et al., 2023). Therefore, we bench-
marked ADD models using EER(%). Lower EER
indicates that the models accurately distinguished
between bonafide and synthetic audios.
Figure 3: Receiver Operating Characteristic (ROC)
Observations: The Aasist and Aasist-L achieve
Curve for Malayalam IndicSynth-IndicSuperb test set
an EER of 0.83% and 0.99% on the LA evaluation created using XTTS-v2. Low Area Under the Curve
set of the ASVspoof 2019 (Jung et al., 2022). Sim- (AUC%) indicates poor discriminative power of ADD
ilarly, the RawNet-2 achieved an EER of 22.38% models.
in the DF track of the ASVspoof 2021 (Liu et al.,
2023). However, as illustrated in Table 2, these synthetic clips. Furthermore, we plotted the Re-
benchmark models achieved significantly higher ceiver Operating Characteristic (ROC) curves as il-
EERs on IndicSynth-IndicSuperb test sets. Ele- lustrated in Figure 3. The ROC curve demonstrates
vated EERs on these test sets indicate that the mod- that the ADD models achieve extremely low Area
els struggled to distinguish between bonafide and Under the Curve (AUC) scores on the Malayalam
4
test set created using the synthetic clips obtained
Aasist: [Link]
RawNet2: [Link] from the XTTS-v2 model. Low AUC indicates
2021/tree/main/DF/Baseline-RawNet2 poor discriminative power of the models. These ob-
22041
servations demonstrate a significant performance Language Source Accuracy ∆ Accuracy (%)
Bonafide 89.925 -
degradation of benchmark ADD models on unseen XTTS-v2 89.763 −0.162
Bengali
language test sets5 . Training such models on large- FreeVC24 90.338 +0.413
scale multilingual synthetic speech datasets can VITS 98.425 +8.500
Bonafide 98.612 -
potentially enhance their generalizability (Müller Gujarati XTTS-v2 96.762 −1.850
et al., 2024b). Therefore, IndicSynth is a valuable FreeVC24 96.475 −2.137
contribution towards generalizable ADD. Bonafide 92.250 -
Hindi XTTS-v2 86.175 −6.075
FreeVC24 85.525 −6.725
4.2 Linguistic Authenticity of IndicSynth Bonafide 88.550 -
Next, we investigate whether IndicSynth’s syn- Kannada XTTS-v2 84.800 −3.750
FreeVC24 85.638 −2.912
thetic speech recordings accurately capture the lin- Bonafide 97.425 -
guistic traits of the target languages. For this ex- Malayalam XTTS-v2 96.200 −1.225
FreeVC24 96.362 −1.063
periment, we created bonafide (IndicSuperb) and
Bonafide 94.900 -
synthetic (IndicSynth) test sets for each genera- Marathi XTTS-v2 89.725 −5.175
tive model and target language, as illustrated in FreeVC24 89.950 −4.950
Bonafide 78.388 -
Table 3. Each test set contains 8,000 audio clips. Punjabi XTTS-v2 66.600 −11.788
These sets contain an equal number of male and FreeVC24 65.938 −12.450
female voice samples. Subsequently, we evaluated Bonafide 41.350 -
Sanskrit XTTS-v2 9.050 −32.300
IndicSynth’s linguistic authenticity by running the FreeVC24 9.175 −32.175
publicly available VoxLingua107 ECAPA-TDNN Bonafide 97.500 -
spoken language identification model (Valk and Tamil XTTS-v2 94.500 −3.000
FreeVC24 94.812 −2.688
Alumäe, 2021; Ravanelli et al., 2021) on these sets6 . Bonafide 98.625 -
The model is trained on the VoxLingua107 dataset, Telugu XTTS-v2 96.100 −2.525
which includes speech recordings from 107 lan- FreeVC24 95.638 −2.987
Bonafide 39.900 -
guages (Valk and Alumae, 2021). Urdu XTTS-v2 33.975 −5.925
Observations: Table 3 illustrates the accuracy FreeVC24 33.763 −6.137
achieved on running the language identification
model through the test sets. We observe an ac- Table 3: Language identification results. We evaluated
IndicSynth’s linguistic authenticity by running language
curacy of more than 80% for most sets. Addi-
identification model through various test sets for each
tionally, we compared the accuracy difference be- generative model and target language (except Odia). We
tween bonafide and synthetic test sets defined as: observe above 80% accuracy in most test sets.
∆ Accuracy%=Accuracysynthetic −Accuracybonafide .
For most languages, the accuracy drop is below ECAPA-TDNN spoken language identification
10%. Interestingly, the accuracies of Bengali syn- model does not support Odia. However, the train-
thetic audios from FreeVC24 and VITS are higher ing set of the model covers 107 languages. There-
than bonafide audios, which indicates that these fore, its embeddings should capture linguistic traits
models are trained on diverse Bengali datasets. effectively. Thus, we obtained the 256-dimensional
language identification embeddings of the Odia
test sets and visualized them through t-SNE using
a perplexity of 40 (Munir et al., 2024), as shown
in Figure 4. Figure 4 shows no clear separation be-
tween the bonafide (IndicSuperb) and the synthetic
(IndicSynth) embeddings. The plot indicates that
IndicSynth’s Odia subset has effectively captured
Figure 4: t-SNE visualization of bonafide (IndicSuperb) linguistic traits of Odia7 .
and synthetic (IndicSynth) Odia dataset. The plot in-
dicates that the IndicSynth-Odia subset has effectively 4.3 Utility of the Mimicry Subset
captured the linguistic traits of Odia. Speaker verification systems accept two speech
Qualitative evaluation: The VoxLingua107 recordings as input and determine if they are from
5
the same speaker. The input speech recordings
Appendix (B) includes additional plots.
6 7
Language identification model: [Link] The t-SNE plots for Punjabi, Sanskrit, and Urdu are in
co/speechbrain/lang-id-voxlingua107-ecapa Appendix (D)

22042
Language SV Model Female Test Set Male Test Set Combined Test Set
ECAPA-TDNN 31.580 29.180 31.110
Bengali ResNet TDNN 23.960 22.960 25.160
X-Vector 43.860 44.520 43.639
ECAPA-TDNN 23.280 34.440 28.880
Gujarati ResNet TDNN 18.520 30.560 24.730
X-Vector 31.040 32.640 32.200
ECAPA-TDNN 25.880 22.700 24.360
Kannada ResNet TDNN 22.740 20.060 21.470
X-Vector 35.020 28.740 31.770
ECAPA-TDNN 30.140 33.680 32.070
Malayalam ResNet TDNN 30.700 31.480 31.170
X-Vector 38.460 38.240 38.520
ECAPA-TDNN 20.620 28.820 24.710
Marathi ResNet TDNN 17.680 25.140 21.600
X-Vector 28.260 31.160 31.080
ECAPA-TDNN 29.300 36.800 32.940
Odia ResNet TDNN 23.580 21.420 25.680
X-Vector 33.440 48.100 42.450
ECAPA-TDNN 25.800 33.420 29.610
Punjabi ResNet TDNN 23.360 30.640 27.050
X-Vector 32.330 36.160 34.350
ECAPA-TDNN 20.160 27.760 23.840
Tamil ResNet TDNN 17.820 25.500 21.540
X-Vector 28.280 31.860 30.070
ECAPA-TDNN 26.380 27.500 26.930
Telugu ResNet TDNN 23.320 25.140 24.550
X-Vector 31.760 32.120 32.180
Table 4: Investigating the vulnerability of state-of-the-art (SOTA) speaker verification models (SV) against imper-
sonation attacks. We observe elevated equal error rates (EER%) when the negative trial pairs contain IndicSynth’s
mimicry subset’s synthetic speech recordings and the target speaker’s bonafide speech from IndicSuperb. It suggests
that the mimicry subset audios closely mimic the bonafide target voices. Therefore, the mimicry subset of IndicSynth
is a valuable resource for enhancing the robustness of SOTA SV models.

form a trial pair. Such systems are vital in forensics, Methodology: We created speaker verification
business, e-commerce, and access control mecha- test sets, as illustrated in Table 4. Each set contains
nisms. However, the malicious use of voice cloning randomly generated 20,000 trial pairs with equal
models may lead to the generation of synthetic positives and negatives. A positive trial pair con-
speech recordings that closely mimic the target tains two bonafide (IndicSuperb) speech recordings
voice. Such synthetic recordings (audio spoofs) of the same target speaker, X. In contrast, a neg-
may be misused to deceive speaker verification ative trial pair contains a bonafide (IndicSuperb)
systems, leading to impersonation attacks against speech recording of a target speaker X and a syn-
the target speaker. Fine-tuning speaker verification thetic (IndicSynth) speech recording of X. Since
models on multilingual synthetic speech datasets each set contains an equal number of male and
can enhance their generalizability and robustness to female speaker trial pairs, we refer to them as com-
out-of-domain audio spoofs. Therefore, this exper- bined test sets. Each combined test set contains
iment explores the utility of IndicSynth’s mimicry 5000 bonafide female trial pairs, 5000 bonafide
subset for enhancing the robustness of speaker veri- male trial pairs, 5000 synthetic female trial pairs,
fication models. We evaluate whether the synthetic and 5000 synthetic male trial pairs. Subsequently,
audios of mimicry subset can deceive three pub- to evaluate gender bias in speaker verification mod-
licly available state-of-the-art (SOTA) speaker ver- els with respect to impersonation attacks, we also
ification models: ECAPA-TDNN, X-Vector, and split the combined test set and created separate
ResNet TDNN (Desplanques et al., 2020; Snyder male and female speaker test sets.
et al., 2018; Villalba et al., 2020)8 . Evaluation Metric: We evaluate the mimicry
8
ECAPA-TDNN:[Link] subset using EER. A higher EER indicates that the
speechbrain/spkrec-ecapa-voxceleb speaker verification model struggled to distinguish
X-Vector:[Link] between positive and negative trial pairs. It im-
spkrec-xvect-voxceleb
ResNet TDNN: [Link] plies that the synthetic speech recordings closely
spkrec-resnet-voxceleb mimic the target speaker’s bonafide voice sample
22043
in a negative trial pair. the t-SNE plots for Odia female and male voices.
Observations: The SOTA speaker verification For each plot, we randomly sampled 500 bonafide
models typically achieve an EER of less than and 500 synthetic clips of the same target speak-
10% when evaluated on unseen language test sets ers. Next, we created t-SNE plots with a perplexity
(Akram et al., 2024; Xia et al., 2019; Mandalapu of 40 using 80-dimensional Mel-Frequency Cep-
et al., 2021). However, as illustrated in Table 4, stral Coefficients (MFCC) features of these audios
we observed significantly elevated EERs ranging (Munir et al., 2024). MFCCs are biologically in-
from 21.470% to 43.639% on the combined test spired speech features that mimic the human au-
sets. Elevated EERs suggest that the speech record- ditory system. The proximity of the bonafide and
ings in IndicSynth’s mimicry subset closely mimic synthetic embeddings in t-SNE indicates that the
the bonafide (IndicSuperb) target voices. Further- mimicry subset’s synthetic audios closely mimic
more, we compared the EERs of the male and fe- the bonafide target voices9 .
male speaker test sets. The EERs of male and fe-
male speaker test sets for the Bengali, Malayalam, 5 Discussion
and Telugu test sets are comparable. However,
This section reflects on our rationale behind cre-
the EER values for Kannada female test sets are
ating mimicry and diversity subsets in IndicSynth.
higher than the male test sets (with absolute differ-
Additionally, we briefly review the potential utility
ences of 2.68% to 6.28% across the speaker veri-
of these subsets with our experimental results.
fication models). This observation indicates that
Rationale behind Mimicry Subset: Indic-
the Kannada female voices mimic the target speak-
Synth’s mimicry subset consists of synthetic voices
ers more closely than the Kannada male voices in
closely mimicking the target speaker’s bonafide
IndicSynth. Similarly, the EERs of the male test
voice. Such synthetic audios—also called au-
sets are higher than the female test sets for Gujarati,
dio spoofs—can deceive speaker verification sys-
Marathi, Odia, Punjabi, and Tamil. This observa-
tems, leading to impersonation attacks. Section
tion indicates that in IndicSynth, male voices in
4.3 demonstrates that speaker verification systems
these languages closely mimic the target speakers
are vulnerable to multilingual audio spoofs (Indic-
compared to female voices.
Synth’s mimicry subset). This observation under-
scores the potential utility of the mimicry subset
for developing multilingual anti-spoofing technolo-
gies.
Need for a Diversity Subset: Synthetic speech
is not only misused for impersonation but also for
spreading misinformation. Synthetic voices cir-
culating in social media to spread misinformation
Figure 5: The t-SNE plot of bonafide (IndicSuperb)
often do not mimic a specific target speaker’s voice.
Odia and IndicSynth’s mimicry subset’s female speak-
ers. The plot reveals the proximity of the bonafide and Instead, misinformation campaigns often involve
synthetic audios. diverse synthetic voices. Therefore, training audio
deepfake detection (ADD) models on a broad range
of synthetic voices is crucial for enhancing the ro-
bustness of these models. However, as we know,
the mimicry subset has a limited speaker diver-
sity as it only includes voices mimicking IndicSu-
perb speakers. Therefore, we introduced the diver-
sity subset in IndicSynth to incorporate a broader
Figure 6: The t-SNE plot of bonafide (IndicSuperb) range of synthetic voices beyond speaker mimicry
Odia and IndicSynth’s mimicry subset’s male speakers. into our dataset. Table 1 provides an overview of
The plot reveals the proximity of the bonafide and syn-
mimicry and diversity subsets.
thetic audios.
Utility of Diversity Subset: As described in Sec-
Qualitative Evaluation: For an extensive evalu-
tion 4.1, we evaluated three state-of-the-art (SOTA)
ation, we visualized the proximity of the bonafide
audio deepfake detection (ADD) models on Indic-
(IndicSuperb) and IndicSynth’s mimicry subset
through t-SNE. Figure 5 and Figure 6 represent 9
Additional t-SNE plots are in Appendix (C).

22044
Synth (both diversity and mimicry subsets) without deepfake detection research.
domain adaptation. These ADD models achieve This work opens up several avenues for research
low Equal Error Rates (EERs) on ASVSpoof chal- on linguistic biases in audio deepfake detection and
lenge datasets. However, we observed a significant anti-spoofing. For instance, IndicSynth can be used
performance degradation of these models on Indic- to investigate the underexplored problem of gender
Synth test sets, as illustrated by elevated EERs and biases in multilingual audio deepfake detection and
low Area Under the Curve (AUC) in Table 2 and for defense against impersonation attacks through
Figure 3. Munir et al. (2024) also reported a similar multilingual spoofs. The dataset is licensed under
observation. In their work, the authors proposed the CC BY-NC 4.0 license.
an Urdu audio deepfake detection dataset. Such
observations indicate a lack of generalizability of 7 Limitations
existing audio deepfake detection models to out-of-
domain languages. It suggests that training or fine- This work introduces IndicSynth, a novel, large-
tuning audio deepfake detection models on multi- scale multilingual synthetic speech dataset to fa-
lingual ADD datasets can potentially enhance the cilitate multlingual audio deepfake detection and
generalizability of these models to out-of-domain anti-spoofing research. We acknowledge that our
languages. Thus, bonafide IndicSuperb data com- work has the following limitations:
bined with the synthetic IndicSynth data (diversity IndicSynth’s Scope and Experimentation: In-
and mimicry subsets) can potentially serve as a dicSynth contains synthetic speech recordings for
valuable dataset for multilingual audio deepfake only 12 languages. Also, the mimicry subsets for
detection. Hindi and Sanskrit are absent in IndicSynth. In
the future, the dataset can be extended by cover-
6 Conclusions and Future Work ing more low-resourced languages and more voice
cloning models for dataset creation. Additionally,
This paper introduces IndicSynth, a novel large- we evaluated IndicSynth using sample test sets. We
scale multilingual synthetic speech dataset to fa- believe that the results from these sets indicate the
cilitate multilingual audio deepfake detection and overall dataset quality.
anti-spoofing research. The dataset contains about Absence of User Study: Ideally, the naturalness
4,000 hours of synthetic audio from 989 target of synthetic speech datasets should be evaluated
speakers, including 456 females and 533 males through a user study. The user study participants
for 12 low-resourced Indian languages. Addition- must be proficient in the target languages for au-
ally, IndicSynth includes rich metadata covering thentic results. However, for large-scale multilin-
the identifiers and gender information of the target gual datasets, such as IndicSynth, recruiting partic-
speakers. Thus facilitating research on gender bi- ipants proficient in low-resource languages is chal-
ases in audio deepfake detection and anti-spoofing. lenging. Thus, meeting our paper’s objectives, we
The dataset consists of mimicry and diversity sub- experimentally evaluated IndicSynth using state-
sets. The mimicry subset includes synthetic au- of-the-art speaker verification models, audio deep-
dios that closely mimic bonafide target voices. In fake detection models, and a language identifica-
contrast, the diversity subset contains a diverse tion model. Additionally, we have included t-SNE
set of realistic synthetic voices. Experimental re- plots in the appendix for a qualitative evaluation of
sults demonstrate that the synthetic audios of the our dataset.
mimicry subset can deceive state-of-the-art (SOTA) We highlight that the challenge of recruiting par-
speaker verification models through impersonation ticipants proficient in low-resourced languages is
attacks. Similarly, empirical results demonstrate not unique to IndicSynth. Instead, it is a common
poor performance of SOTA audio deepfake detec- issue faced by researchers working towards gener-
tion models on Indian language test sets. Further- ating multilingual datasets for social good. How-
more, qualitative and quantitative evaluation using ever, with around 7,000 global languages and rising
a SOTA language identification model validated cases of deepfake-related fraud, there is an urgent
the linguistic authenticity of our dataset. It turns need for multilingual synthetic datasets to facilitate
out that IndicSynth is a valuable contribution to research on multilingual audio deepfake detection.
preventing impersonation attacks on speaker verifi- Therefore, the absence of user studies should not
cation systems and facilitating multilingual audio hinder the release of such datasets. Instead, the
22045
community members with access to computational Jordan J. Bird and Ahmad Lotfi. 2023. Real-time de-
resources can contribute by constructing and releas- tection of ai-generated speech for deepfake voice
conversion. ArXiv, abs/2308.12734.
ing more multilingual datasets. Subsequently, the
members who can connect to native speakers of Brecht Desplanques, Jenthe Thienpondt, and Kris De-
those languages can contribute by conducting user muynck. 2020. ECAPA-TDNN: emphasized chan-
studies to evaluate human perception of synthetic nel attention, propagation and aggregation in TDNN
based speaker verification. In Interspeech 2020,
speech. pages 3830–3834. ISCA.
Despite these limitations, the dearth of multilin-
gual synthetic speech datasets makes IndicSynth a Gölge Eren and The Coqui TTS Team. 2021. Coqui
valuable resource that can facilitate research on TTS.
multilingual audio deepfake detection and anti- Joel Cameron Frank and Lea Schönherr. 2021. Wave-
spoofing. fake: A data set to facilitate audio deepfake detection.
ArXiv, abs/2111.02813.
8 Ethical Considerations Kurtis Haut, Caleb Wohn, Victor Antony, Aidan
Goldfarb, Melissa Welsh, Dillanie Sumanthiran,
Synthetic speech datasets are essential to advance M. Rafayet Ali, and Ehsan Hoque. 2022. Demo-
audio deepfake detection and anti-spoofing re- graphic feature isolation for bias research using deep-
search. However, we realize that such datasets can fakes. In Proceedings of the 30th ACM Interna-
inadvertently contribute to the refinement of au- tional Conference on Multimedia, MM ’22, page
6890–6897, New York, NY, USA. Association for
dio deepfake generation technologies by malicious Computing Machinery.
users. Therefore, responsible management of these
resources is crucial. Thus, we release IndicSynth Tahir Javed, Kaushal Bhogale, Abhigyan Raman,
under CC BY-NC 4.0, restricting commercial use Pratyush Kumar, Anoop Kunchukuttan, and Mitesh
Khapra. 2023. Indicsuperb: A speech processing uni-
of our dataset. Furthermore, IndicSynth is a syn- versal performance benchmark for indian languages.
thetic speech dataset generated from the publicly Proceedings of the AAAI Conference on Artificial
available bonafide IndicSuperb dataset. IndicSu- Intelligence, 37:12942–12950.
perb is licensed under the Creative Commons CC0
Yan Ju, Shu Hu, Shan Jia, George H. Chen, and Si-
license (“no rights reserved”). The CC0 license wei Lyu. 2024. Improving Fairness in Deepfake
allows users to freely build upon, reuse, or enhance Detection . In 2024 IEEE/CVF Winter Conference
the dataset without restriction. We strongly encour- on Applications of Computer Vision (WACV), pages
age the community to use IndicSynth for social 4643–4653, Los Alamitos, CA, USA. IEEE Com-
puter Society.
good and advance research on multilingual audio
deepfake detection and anti-spoofing. Jee-weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-jin
Shim, Joon Son Chung, Bong-Jin Lee, Ha-Jin Yu, and
Acknowledgments Nicholas Evans. 2022. Aasist: Audio anti-spoofing
using integrated spectro-temporal graph attention net-
This work is supported by the Infosys Centre for works. In ICASSP 2022 - 2022 IEEE International
Conference on Acoustics, Speech and Signal Process-
Artificial Intelligence (CAI) at IIIT-Delhi. We also ing (ICASSP), pages 6367–6371.
thank the SBILab at IIIT-Delhi for their helpful
discussions and support. Piotr Kawa, Marcin Plata, and Piotr Syga. 2022. Attack
agnostic dataset: Towards generalization and stabi-
lization of audio deepfake detection. pages 4023–
4027.
References
Hasam Khalid, Shahroz Tariq, and Simon S. Woo. 2021.
Ali Akram, Marija Stanojevic, Malikeh Ehghaghi, Fakeavceleb: A novel audio-video multimodal deep-
and Jekaterina Novikova. 2024. Zero-shot multi- fake dataset. ArXiv, abs/2108.05080.
lingual speaker verification in clinical trials. ArXiv,
abs/2404.01981. Pavel Korshunov and Sébastien Marcel. 2022. Im-
proving generalization of deepfake detection with
Zhongjie Ba, Qing Wen, Peng Cheng, Yuwei Wang, data farming and few-shot learning. IEEE Transac-
Feng Lin, Li Lu, and Zhenguang Liu. 2023. Trans- tions on Biometrics, Behavior, and Identity Science,
ferring audio deepfake detection capability across 4(3):386–397.
languages. In Proceedings of the ACM Web Confer-
ence 2023, WWW ’23, page 2033–2044, New York, Xuechen Liu, Xin Wang, Md Sahidullah, Jose Patino,
NY, USA. Association for Computing Machinery. Héctor Delgado, Tomi Kinnunen, Massimiliano
22046
Todisco, Junichi Yamagishi, Nicholas Evans, An- In Findings of the Association for Computational Lin-
dreas Nautsch, and Kong Aik Lee. 2023. Asvspoof guistics: NAACL 2024, pages 379–394, Mexico City,
2021: Towards spoofed and deepfake speech detec- Mexico. Association for Computational Linguistics.
tion in the wild. IEEE/ACM Transactions on Audio,
Speech, and Language Processing, 31:2507–2522. Divya Sharma and Arun Balaji Buduru. 2022. FAt-
Net: Cost-effective approach towards mitigating the
Hareesh Mandalapu, Thomas Møller Elbo, Raghavendra linguistic bias in speaker verification systems. In
Ramachandra, and Christoph Busch. 2021. Cross- Findings of the Association for Computational Lin-
lingual speaker verification: Evaluation on x-vector guistics: NAACL 2022, pages 1247–1258, Seattle,
method. In Intelligent Technologies and Applica- United States. Association for Computational Lin-
tions, pages 215–226, Cham. Springer International guistics.
Publishing.
David Snyder, Daniel Garcia-Romero, Alan McCree,
Sheza Munir, Wassay Sajjad, Mukeet Raza, Emaan Ab- Gregory Sell, Daniel Povey, and Sanjeev Khudanpur.
bas, Abdul Hameed Azeemi, Ihsan Ayyub Qazi, and 2018. Spoken language recognition using x-vectors.
Agha Ali Raza. 2024. Deepfake defense: Construct- In The Speaker and Language Recognition Workshop
ing and evaluating a specialized Urdu deepfake audio (Odyssey 2018), pages 105–111.
dataset. In Findings of the Association for Computa-
tional Linguistics: ACL 2024, pages 14470–14480, Sidharth T and Guhan T. 2024. Deepfake technology in
Bangkok, Thailand. Association for Computational social media: Social and legal implications in india.
Linguistics. IJFMR, 6(6).

Nicolas Müller, Nicholas Evans, Hemlata Tak, Philip Hemlata Tak, Jose Patino, Massimiliano Todisco, An-
Sperl, and Konstantin Böttinger. 2024a. Harder or dreas Nautsch, Nicholas Evans, and Anthony Larcher.
different? understanding generalization of audio 2021. End-to-end anti-spoofing with rawnet2. In
deepfake detection. pages 2705–2709. ICASSP 2021 - 2021 IEEE International Confer-
ence on Acoustics, Speech and Signal Processing
Nicolas M. Müller, Piotr Kawa, Wei Herng Choong, (ICASSP), pages 6369–6373.
Edresson Casanova, Eren Gölge, Thorsten Müller,
Piotr Syga, Philip Sperl, and Konstantin Böttinger. Jörgen Valk and Tanel Alumäe. 2021. VoxLingua107:
2024b. Mlaad: The multi-language audio anti- a dataset for spoken language recognition. In Proc.
spoofing dataset. In 2024 International Joint Confer- IEEE SLT Workshop.
ence on Neural Networks (IJCNN), pages 1–7.
Jorgen Valk and Tanel Alumae. 2021. Voxlingua107:
Xiaoke Qi, Hao Gu, Jiangyan Yi, Jianhua Tao, Yong A dataset for spoken language recognition. pages
Ren, Jiayi He, and Siding Zeng. 2024. Madd: A 652–658.
multi-lingual multi-speaker audio deepfake detection
dataset. In 2024 IEEE 14th International Symposium Jesús Villalba, Nanxin Chen, David Snyder, Daniel
on Chinese Spoken Language Processing (ISCSLP), Garcia-Romero, Alan McCree, Gregory Sell,
pages 466–470. Jonas Borgstrom, Leibny Paola García-Perera,
Fred Richardson, Réda Dehak, Pedro A. Torres-
Mouna Rabhi, Spiridon Bakiras, and Roberto Di Pietro. Carrasquillo, and Najim Dehak. 2020. State-of-the-
2024. Audio-deepfake detection: Adversarial attacks art speaker recognition with neural network embed-
and countermeasures. Expert Systems with Applica- dings in nist sre18 and speakers in the wild evalua-
tions, 250:123941. tions. Computer Speech & Language, 60:101026.

Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Xin Wang, Junichi Yamagishi, Massimiliano Todisco,
Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Héctor Delgado, Andreas Nautsch, Nicholas Evans,
Subakan, Nauman Dawalatabad, Abdelwahab Heba, Md Sahidullah, Ville Vestman, Tomi Kinnunen,
Jianyuan Zhong, Ju-Chieh Chou, Sung-Lin Yeh, Kong Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai
Szu-Wei Fu, Chien-Feng Liao, Elena Rastorgueva, Peng, Hsin-Te Hwang, Yu Tsao, Hsin-Min Wang,
François Grondin, William Aris, Hwidong Na, Yan Sébastien Le Maguer, Markus Becker, Fergus Hen-
Gao, Renato De Mori, and Yoshua Bengio. 2021. derson, Rob Clark, Yu Zhang, Quan Wang, Ye Jia,
SpeechBrain: A general-purpose speech toolkit. Kai Onuma, Koji Mushika, Takashi Kaneda, Yuan
Preprint, arXiv:2106.04624. ArXiv:2106.04624. Jiang, Li-Juan Liu, Yi-Chiao Wu, Wen-Chin Huang,
Tomoki Toda, Kou Tanaka, Hirokazu Kameoka, In-
ARTH JUHUL SHAH, Ravindrakumar M. Purohit, gmar Steiner, Driss Matrouf, Jean-François Bonas-
Dharmendra H. Vaghera, and Hemant Patil. 2024. tre, Avashna Govender, Srikanth Ronanki, Jing-Xuan
MLADDC: Multi-lingual audio deepfake detection Zhang, and Zhen-Hua Ling. 2020. Asvspoof 2019: A
corpus. In Audio Imagination: NeurIPS 2024 Work- large-scale public database of synthesized, converted
shop AI-Driven Speech, Music, and Sound Genera- and replayed speech. Computer Speech & Language,
tion. 64:101114.

Divya Sharma. 2024. EcoSpeak: Cost-efficient bias mit- Wei Xia, Jing Huang, and John H.L. Hansen. 2019.
igation for partially cross-lingual speaker verification. Cross-lingual text-independent speaker verification
22047
IndicSynth/
using unsupervised adversarial discriminative do- |-- <language>/
main adaptation. In ICASSP 2019 - 2019 IEEE Inter- | |-- XTTS_v2/
national Conference on Acoustics, Speech and Signal | | |-- Male/
Processing (ICASSP), pages 5816–5820. | | | |-- <speaker_id>/
| | | | |-- <clip_id>.wav
Yuankun Xie, Haonan Cheng, Yutian Wang, and Long | | | | `-- ...
Ye. 2024. Domain generalization via aggregation | | |-- Female/
and separation for audio deepfake detection. IEEE | | | |-- <speaker_id>/
Transactions on Information Forensics and Security, | | | | |-- <clip_id>.wav
19:344–358. | | | | `-- ...
| | `-- [Link]
Ying Xu, Philipp Terhörst, Marius Pedersen, and Kiran | |-- FreeVC24/
Raja. 2024. Analyzing fairness in deepfake detection | | |-- Male/
with massively annotated databases. IEEE Transac- | | | |-- <speaker_id>/
tions on Technology and Society, 5(1):93–106. | | | | |-- <clip_id>.wav
| | | | `-- ...
Jiangyan Yi, Ruibo Fu, Jianhua Tao, Shuai Nie, Haoxin | | |-- Female/
Ma, Chenglong Wang, Tao Wang, Zhengkun Tian, | | | |-- <speaker_id>/
Ye Bai, Cunhang Fan, Shan Liang, Shiming Wang, | | | | |-- <clip_id>.wav
| | | | `-- ...
Shuai Zhang, Xinrui Yan, Le Xu, Zhengqi Wen, and | | `-- [Link]
Haizhou Li. 2022. Add 2022: the first audio deep
synthesis detection challenge. In ICASSP 2022 -
2022 IEEE International Conference on Acoustics, Figure 7: IndicSynth directory structure.
Speech and Signal Processing (ICASSP), pages 9216–
9220.
B IndicSynth for Audio DeepFake
Jiangyan Yi, Jianhua Tao, Ruibo Fu, Xinrui Yan, Chen- Detection: Additional Plots
glong Wang, Tao Wang, Chu Yuan Zhang, Xiao-
hui Zhang, Yan Zhao, Yong Ren, Leling Xu, Jun Figures 8–56 illustrate the Receiver Operating
Zhou, Hao Gu, Zhengqi Wen, Shan Liang, Zheng
Lian, Shuai Nie, and Haizhou Li. 2023. Add 2023:
Characteristic (ROC) Curves and the Detection Er-
the second audio deepfake detection challenge. In ror Trade-off (DET) Curves for IndicSynth test sets
DADA@IJCAI. created using various generative models. The plots
indicate that the benchmark audio deepfake detec-
Jiangyan Yi, Chu Yuan Zhang, Jianhua Tao, Cheng-
long Wang, Xinrui Yan, Yong Ren, Hao Gu, and tion models lack generalizability on out-of-domain
Junzuo Zhou. 2024. Add 2023: Towards audio deep- (Indian) language test sets.
fake detection and analysis in the wild. Preprint,
arXiv:2408.04967.
Mohammed Yousif, Jonat John Mathew, Huzaifa Pal-
lan, Agamjeet Singh Padda, Syed Daniyal Shah, Sara
Adamski, Madhu Reddiboina, and Arjun Pankajak-
shan. 2024. Enhancing generalization in audio deep-
fake detection: A neural collapse based sampling and
training approach. Preprint, arXiv:2404.13008.
Yi Zhu, Surya Koppisetti, Trang Tran, and Gaurav
Bharaj. 2024. Slim: Style-linguistics mismatch
model for generalized audio deepfake detection.
ArXiv, abs/2407.18517.
Figure 8: The Receiver Operating Characteristic (ROC)
A IndicSynth Directory Structure Curve for the Bengali test set created using XTTS-v2.
Low Area Under the Curve (AUC) values of audio
Figure 7 shows the directory structure of Indic- deepfake detection models indicate poor discriminative
Synth. The dataset contains a folder for each target power of these models on our test set.
language. Each language folder includes subfold-
ers for generative models: XTTS_v2, FreeVC24,
or the VITS. Each generative model folder has a
metadata file and subfolders for anonymous target
speaker IDs. Within each target speaker’s folder,
there are synthetic audio clips. The speaker IDs
used in the IndicSynth are the same as those used
in IndicSuperb.
22048
Figure 9: The Detection Error Trade-off (DET) Curve Figure 12: The Receiver Operating Characteristic
for the Bengali test set created using XTTS-v2. The (ROC) Curve for the Bengali test set created using
trend of DET curves towards the upper right indicates VITS. Low Area Under the Curve (AUC) values of
the poor capability of the audio deepfake detection mod- audio deepfake detection models indicate poor discrimi-
els in distinguishing between bonafide and synthetic native power of these models on our test set.
audios.

Figure 10: The Receiver Operating Characteristic Figure 13: The Detection Error Trade-off (DET) Curve
(ROC) Curve for the Bengali test set created using for the Bengali test set created using VITS. The trend of
FreeVC24. Low Area Under the Curve (AUC) values DET curves towards the upper right indicates the poor
of audio deepfake detection models indicate poor dis- capability of the audio deepfake detection models in
criminative power of these models on our test set. distinguishing between bonafide and synthetic audios.

Figure 11: The Detection Error Trade-off (DET) Curve


for the Bengali test set created using FreeVC24. The Figure 14: The Receiver Operating Characteristic
trend of DET curves towards the upper right indicates (ROC) Curve for the Gujarati test set created using
the poor capability of the audio deepfake detection mod- XTTS-v2. Low Area Under the Curve (AUC) values of
els in distinguishing between bonafide and synthetic audio deepfake detection models indicate poor discrimi-
audios. native power of these models on our test set.

22049
Figure 15: The Detection Error Trade-off (DET) Curve
for the Gujarati test set created using XTTS-v2. The Figure 18: The Receiver Operating Characteristic
trend of DET curves towards the upper right indicates (ROC) Curve for the Hindi test set created using XTTS-
the poor capability of the audio deepfake detection mod- v2. Low Area Under the Curve (AUC) values of audio
els in distinguishing between bonafide and synthetic deepfake detection models indicate poor discriminative
audios. power of these models on our test set.

Figure 16: The Receiver Operating Characteristic


(ROC) Curve for the Gujarati test set created using
Figure 19: The Detection Error Trade-off (DET) Curve
FreeVC24. Low Area Under the Curve (AUC) values
for the Hindi test set created using XTTS-v2.
of audio deepfake detection models indicate poor dis-
criminative power of these models on our test set.

Figure 17: The Detection Error Trade-off (DET) Curve Figure 20: Receiver Operating Characteristic (ROC)
for the Gujarati test set created using FreeVC24. The Curve for the Hindi IndicSynth-IndicSuperb test set
trend of DET curves towards the upper right indicates created using FreeVC24. Low Area Under the Curve
the poor capability of the audio deepfake detection mod- (AUC%) indicates poor discriminative power of ADD
els in distinguishing between bonafide and synthetic models.
audios.

22050
Figure 21: The Detection Error Trade-off (DET) Curve Figure 24: The Receiver Operating Characteristic
for the Hindi test set created using FreeVC24. The trend (ROC) Curve for the Kannada test set created using
of DET curves towards the upper right indicates the FreeVC24. Low Area Under the Curve (AUC) values
poor capability of the audio deepfake detection models of audio deepfake detection models indicate poor dis-
in distinguishing between bonafide and synthetic audios. criminative power of these models on our test set.

Figure 25: The Detection Error Trade-off (DET) Curve


Figure 22: The Receiver Operating Characteristic
for the Kannada test set created using FreeVC24. The
(ROC) Curve for the Kannada test set created using
trend of DET curves towards the upper right indicates
XTTS-v2. Low Area Under the Curve (AUC) values of
the poor capability of the audio deepfake detection mod-
audio deepfake detection models indicate poor discrimi-
els in distinguishing between bonafide and synthetic
native power of these models on our test set.
audios.

Figure 23: The Detection Error Trade-off (DET) Curve Figure 26: The Detection Error Trade-off (DET) Curve
for the Kannada test set created using XTTS-v2. The for the Malayalam test set created using XTTS-v2. The
trend of DET curves towards the upper right indicates trend of DET curves towards the upper right indicates
the poor capability of the audio deepfake detection mod- the poor capability of the audio deepfake detection mod-
els in distinguishing between bonafide and synthetic els in distinguishing between bonafide and synthetic
audios. audios.

22051
Figure 27: The Receiver Operating Characteristic Figure 30: The Detection Error Trade-off (DET) Curve
(ROC) Curve for the Malayalam test set created using for the Marathi test set created using XTTS-v2. The
FreeVC24. Low Area Under the Curve (AUC) values trend of DET curves towards the upper right indicates
of audio deepfake detection models indicate poor dis- the poor capability of the audio deepfake detection mod-
criminative power of these models on our test set. els in distinguishing between bonafide and synthetic
audios.

Figure 28: The Detection Error Trade-off (DET) Curve


Figure 31: The Receiver Operating Characteristic
for the Malayalam test set created using FreeVC24. The
(ROC) Curve for the Marathi test set created using
trend of DET curves towards the upper right indicates
FreeVC24. Low Area Under the Curve (AUC) values
the poor capability of the audio deepfake detection mod-
of audio deepfake detection models indicate poor dis-
els in distinguishing between bonafide and synthetic
criminative power of these models on our test set.
audios.

Figure 32: The Detection Error Trade-off (DET) Curve


Figure 29: The Receiver Operating Characteristic for the Marathi test set created using FreeVC24. The
(ROC) Curve for the Marathi test set created using trend of DET curves towards the upper right indicates
XTTS-v2. Low Area Under the Curve (AUC) values of the poor capability of the audio deepfake detection mod-
audio deepfake detection models indicate poor discrimi- els in distinguishing between bonafide and synthetic
native power of these models on our test set. audios.

22052
Figure 33: The Receiver Operating Characteristic Figure 36: The Detection Error Trade-off (DET) Curve
(ROC) Curve for the Odia test set created using XTTS- for the Odia test set created using FreeVC24. The trend
v2. Low Area Under the Curve (AUC) values of audio of DET curves towards the upper right indicates the
deepfake detection models indicate poor discriminative poor capability of the audio deepfake detection models
power of these models on our test set. in distinguishing between bonafide and synthetic audios.

Figure 37: The Receiver Operating Characteristic


Figure 34: The Detection Error Trade-off (DET) Curve
(ROC) Curve for the Punjabi test set created using
for the Odia test set created using XTTS-v2. The trend
XTTS-v2. Low Area Under the Curve (AUC) values of
of DET curves towards the upper right indicates the
audio deepfake detection models indicate poor discrimi-
poor capability of the audio deepfake detection models
native power of these models on our test set.
in distinguishing between bonafide and synthetic audios.

Figure 38: The Detection Error Trade-off (DET) Curve


Figure 35: The Receiver Operating Characteristic for the Punjabi test set created using XTTS-v2. The
(ROC) Curve for the Odia test set created using trend of DET curves towards the upper right indicates
FreeVC24. Low Area Under the Curve (AUC) values the poor capability of the audio deepfake detection mod-
of audio deepfake detection models indicate poor dis- els in distinguishing between bonafide and synthetic
criminative power of these models on our test set. audios.

22053
Figure 39: The Receiver Operating Characteristic Figure 42: The Detection Error Trade-off (DET) Curve
(ROC) Curve for the Punjabi test set created using for the Sanskrit test set created using XTTS-v2. The
FreeVC24. Low Area Under the Curve (AUC) values trend of DET curves towards the bottom left indicates
of audio deepfake detection models indicate poor dis- the capability of the audio deepfake detection models in
criminative power of these models on our test set. distinguishing between bonafide and synthetic audios.

Figure 43: The Receiver Operating Characteristic


Figure 40: The Detection Error Trade-off (DET) Curve (ROC) Curve for the Sanskrit test set created using
for the Punjabi test set created using FreeVC24. The FreeVC24. Low Area Under the Curve (AUC) values
trend of DET curves towards the upper right indicates of audio deepfake detection models indicate poor dis-
the poor capability of the audio deepfake detection mod- criminative power of these models on our test set.
els in distinguishing between bonafide and synthetic
audios.

Figure 44: The Detection Error Trade-off (DET) Curve


for the Sanskrit test set created using FreeVC24. The
trend of DET curves towards the upper right indicates
Figure 41: The Receiver Operating Characteristic the poor capability of the audio deepfake detection mod-
(ROC) Curve for the Sanskrit test set created using els in distinguishing between bonafide and synthetic
XTTS-v2. audios.

22054
Figure 45: The Receiver Operating Characteristic
(ROC) Curve for the Tamil test set created using XTTS-
v2. Low Area Under the Curve (AUC) values of audio Figure 47: The Receiver Operating Characteristic
deepfake detection models indicate poor discriminative (ROC) Curve for the Tamil test set created using
power of these models on our test set. FreeVC24. Low Area Under the Curve (AUC) values
of audio deepfake detection models indicate poor dis-
criminative power of these models on our test set.

Figure 46: The Detection Error Trade-off (DET) Curve


for the Tamil test set created using XTTS-v2. The trend
of DET curves towards the upper right indicates the
poor capability of the audio deepfake detection models
in distinguishing between bonafide and synthetic audios.
Figure 48: The Detection Error Trade-off (DET) Curve
for the Tamil test set created using FreeVC24. The trend
of DET curves towards the upper right indicates the poor
capability of the audio deepfake detection models in
distinguishing between bonafide and synthetic audios.

Figure 49: The Receiver Operating Characteristic


(ROC) Curve for the Telugu test set created using XTTS-
v2. Low Area Under the Curve (AUC) values of audio
deepfake detection models indicate poor discriminative
power of these models on our test set.

22055
Figure 50: The Detection Error Trade-off (DET) Curve Figure 53: The Receiver Operating Characteristic
for the Telugu test set created using XTTS-v2. The trend (ROC) Curve for the Urdu test set created using XTTS-
of DET curves towards the upper right indicates the poor v2. Low Area Under the Curve (AUC) values of audio
capability of the audio deepfake detection models in deepfake detection models indicate poor discriminative
distinguishing between bonafide and synthetic audios. power of these models on our test set.

Figure 51: The Receiver Operating Characteristic


(ROC) Curve for the Telugu test set created using Figure 54: The Detection Error Trade-off (DET) Curve
FreeVC24. Low Area Under the Curve (AUC) values for the Urdu test set created using XTTS-v2. The trend
of audio deepfake detection models indicate poor dis- of DET curves towards the upper right indicates the
criminative power of these models on our test set. poor capability of the audio deepfake detection models
in distinguishing between bonafide and synthetic audios.

Figure 52: The Detection Error Trade-off (DET) Curve


for the Telugu test set created using FreeVC24. The Figure 55: The Receiver Operating Characteristic
trend of DET curves towards the upper right indicates (ROC) Curve for the Urdu test set created using
the poor capability of the audio deepfake detection mod- FreeVC24. Low Area Under the Curve (AUC) values
els in distinguishing between bonafide and synthetic of audio deepfake detection models indicate poor dis-
audios. criminative power of these models on our test set.

22056
Figure 58: The t-SNE plot of bonafide (IndicSuperb)
Figure 56: The Detection Error Trade-off (DET) Curve Gujarati and IndicSynth’s mimicry subset’s male speak-
for the Urdu test set created using FreeVC24. The trend ers. The plot reveals the proximity of the bonafide and
of DET curves towards the upper right indicates the synthetic audios.
poor capability of the audio deepfake detection models
in distinguishing between bonafide and synthetic audios.

C Authenticity of Mimicry Subset


Figures 57–73 illustrate the proximity of the
mimicry subset audios with the bonafide IndicSu-
perb audios for the target languages.

Figure 59: The t-SNE plot of bonafide (IndicSuperb)


Kannada and IndicSynth’s mimicry subset’s female
speakers. The plot reveals the proximity of the bonafide
and synthetic audios.

Figure 57: The t-SNE plot of bonafide (IndicSuperb)


Gujarati and IndicSynth’s mimicry subset’s female
speakers. The plot reveals the proximity of the bonafide
and synthetic audios. Figure 60: The t-SNE plot of bonafide (IndicSuperb)
Kannada and IndicSynth’s mimicry subset’s male speak-
ers. The plot reveals the proximity of the bonafide and
synthetic audios.

Figure 61: The t-SNE plot of bonafide (IndicSuperb)


Malayalam and IndicSynth’s mimicry subset’s female
speakers. The plot reveals the proximity of the bonafide
and synthetic audios.
22057
Figure 66: The t-SNE plot of bonafide (IndicSuperb)
Punjabi and IndicSynth’s mimicry subset’s male speak-
ers. The plot reveals the proximity of the bonafide and
synthetic audios.
Figure 62: The t-SNE plot of bonafide (IndicSuperb)
Malayalam and IndicSynth’s mimicry subset’s male
speakers. The plot reveals the proximity of the bonafide
and synthetic audios.

Figure 67: The t-SNE plot of bonafide (IndicSuperb)


Tamil voices and IndicSynth’s mimicry subset’s syn-
thetic clips of the same target female speakers. The
plot reveals the proximity of the bonafide and synthetic
Figure 63: The t-SNE plot of bonafide (IndicSuperb) audios.
Marathi and IndicSynth’s mimicry subset’s female
speakers. The plot reveals the proximity of the bonafide
and synthetic audios.

Figure 68: The t-SNE plot of bonafide (IndicSuperb)


Figure 64: The t-SNE plot of bonafide (IndicSuperb)
Tamil voices and IndicSynth’s mimicry subset’s syn-
Marathi and IndicSynth’s mimicry subset’s male speak-
thetic clips of the same target male speakers. The plot
ers. The plot reveals the proximity of the bonafide and
reveals the proximity of the bonafide and synthetic au-
synthetic audios.
dios.

Figure 69: The t-SNE plot of bonafide (IndicSuperb)


Tamil voices and IndicSynth’s mimicry subset’s syn-
thetic clips of the same target female speakers. The
Figure 65: The t-SNE plot of bonafide (IndicSuperb) plot reveals the proximity of the bonafide and synthetic
Punjabi and IndicSynth’s mimicry subset’s female audios.
speakers. The plot reveals the proximity of the bonafide
and synthetic audios.
22058
D IndicSynth Linguistic Authenticity:
Additional Plots
Figures 74–76 illustrate the t-SNE plots from the
embeddings obtained through the language iden-
tification model for Sanskrit, Punjabi, and Urdu.
The plots demonstrate the linguistic authenticity of
IndicSynth’s Sanskrit, Punjabi, and Urdu audios.

Figure 70: The t-SNE plot of bonafide (IndicSuperb)


Telugu and IndicSynth’s mimicry subset’s female speak-
ers. The plot reveals the proximity of the bonafide and
synthetic audios.

Figure 74: t-SNE visualization of bonafide (IndicSu-


perb) and synthetic (IndicSynth) Sanskrit dataset. The
plot indicates that the IndicSynth-Sanskrit subset has
effectively captured Sanskrit’s linguistic traits.

Figure 71: The t-SNE plot of bonafide (IndicSuperb)


Telugu and IndicSynth’s mimicry subset’s male speak-
ers. The plot reveals the proximity of the bonafide and
synthetic audios.
Figure 75: t-SNE visualization of bonafide (IndicSu-
perb) and synthetic (IndicSynth) Punjabi dataset. The
plot indicates that the IndicSynth-Punjabi subset has
effectively captured Punjabi’s linguistic traits.

Figure 72: The t-SNE plot of bonafide (IndicSuperb)


Urdu and IndicSynth’s mimicry subset’s female speak- Figure 76: t-SNE visualization of bonafide (IndicSu-
ers. The plot reveals the proximity of the bonafide and perb) and synthetic (IndicSynth) Urdu dataset. The plot
synthetic audios. indicates that the IndicSynth-Urdu subset has effectively
captured Urdu’s linguistic traits.

E Costs
This section highlights the costs of various experi-
ments conducted for this study in terms of carbon
emissions, electricity consumption, and execution
time. The data generation and experimentation
were done using the NVIDIA RTX A6000 GPU.
Below are the details:
Figure 73: The t-SNE plot of bonafide (IndicSuperb)
Urdu and IndicSynth’s mimicry subset’s male speakers. 1. Cost of fine-tuning XTTS-v2: Fine-tuning
The plot reveals the proximity of the bonafide and syn-
the XTTS-v2 for one epoch for a particular
thetic audios.
target language takes approximately 1.5 hours.
22059
In those 1.5 hours, the process causes approx- 22,150,912 parameters and occupies 364 MB
imately 0.403 kgCO2eq carbon emission and of memory.
consumes 0.564 kWh of electricity. This re- 8. X-Vector speaker verification model: It took
sult is for the training set containing 70,692 about 1.5 minutes to generate ResNet TDNN
audio clips. We fine-tuned each model for embeddings for 10,000 audio clips. The
45 epochs. The XTTS-v2 model contains model caused approximately 0.003 kgCO2eq
470,751,571 parameters. carbon emission and energy consumption of
2. XTTS-v2: It takes about 45 minutes to gen- 0.004 kWh for 10,000 embeddings. The X-
erate 1,000 audio clips from the XTTS-v2 Vector model has 8,172,473 parameters and
model for Hindi. The average duration of a occupies 300 MB of memory.
Hindi audio clip in IndicSuperb is 2.65 sec- 9. AASIST Audio Deepfake Detection Model:
onds. For generating 1000 synthetic audios us- The AASIST model has 297,866 parameters
ing the XTTS-v2 model, approximately 0.106 and occupies 264 MB of memory.
kgCO2eq carbon emission and 0.148 kWh 10. AASIST-L Audio Deepfake Detection
of electricity are consumed. The XTTS-v2 Model: The AASIST model has 85,306 pa-
model has 470,751,571 parameters and occu- rameters and occupies 262 MB of memory.
pies 2,110 MB of memory. 11. RawNet-2 Audio DeepFake Detection
3. FreeVC24: It takes about 4 minutes to gen- Model: The RawNet-2 model contains
erate 1000 audio clips from the FreeVC24 17,623,671 parameters and occupies 67.2 MB
model for Hindi. The average duration of of disk space.
a Hindi audio clip in IndicSuperb is 2.65
seconds. For generating 1000 synthetic au- F Tools and Software used
dios using the FreeVC24 model, approxi- We used the following tools and software for this
mately 0.010 kgCO2eq carbon emissions and study (other than the ones already cited in the pa-
0.014 kWh of electricity are consumed. The per):
FreeVC24 model contains 356,216,448 pa-
rameters and occupies 1690 MB of memory. 1. We used Grammarly and ChatGPT for better
4. VITS: The VITS model has 83,050,540 pa- sentence construction at occasional places and
rameters and occupies 586 MB of memory. to enhance clarity in our draft.
5. Language identification: It took about 2 min- 2. We used [Link] and the matplotlib for dia-
utes to run the VoxLingua107 ECAPA-TDNN grams.
spoken language identification model on a test 3. We used Librosa to generate the MFCC fea-
set containing 8,000 audio clips. The model tures.
caused approximately 0.004 kgCO2eq carbon 4. We used Pytorch for experimentation: version
emission and energy consumption of 0.006 2.5.1+cu124.
kWh.
6. ResNet TDNN speaker verification model: G Licenses
It took about 4 minutes to generate ResNet In this section, we specify the licenses of the
TDNN embeddings for 10,000 audio clips. datasets and models that we have used for this
The model caused approximately 0.012 study.
kgCO2eq carbon emission and energy con-
sumption of 0.018 kWh for 10,000 em- 1. IndicSuperb: Creative Commons CC0 license
beddings. The ResNet TDNN model has ("no rights reserved").
17,282,816 parameters and occupies 334 MB 2. Aasist, Aasist-L, and the RawNet-2: The
of memory. ADD models used are licensed under the MIT
7. Ecapa-TDNN speaker verification model: License.
It took about 3 minutes to generate ResNet 3. The speechbrain models (Ecapa-TDNN,
TDNN embeddings for 10,000 audio clips. ResNet TDNN, X-Vector, VoxLingua107
The model caused approximately 0.007 ECAPA-TDNN spoken language identifica-
kgCO2eq carbon emission and energy con- tion model) is licensed under the Apache Li-
sumption of 0.009 kWh for 10,000 em- cense 2.0.
beddings. The Ecapa-TDNN model has
22060

You might also like