Indicsynth: A Large-Scale Multilingual Synthetic Speech Dataset For Low-Resource Indian Languages
Indicsynth: A Large-Scale Multilingual Synthetic Speech Dataset For Low-Resource Indian Languages
Males 446.9 464.3 The generated synthetic speech (v tgt tts ) articu-
400
lates the transcript (ttxt ) while attempting to mimic
300 270.9
209.7 207.3 218.6 the target voice (v tgt ). In contrast to the TTS mod-
200 169.0 172.1 185.1
170.0
96.2 86.4110.5
170.1 els, voice conversion (VC) models take a bonafide
100 78.3 84.6 64.981.3
55.2 58.6 58.236.8
13.3 source speech recording (v src ) and a bonafide tar-
0
get speech recording (v tgt ) as inputs. Subsequently,
Gu ali
ati
Ka di
lay a
Ma m
hi
Pu ia
Sa abi
t
Tel l
u
u
i
kri
Tam
Ma nad
ug
Urd
Od
Hin
rat
ala
ng
jar
nj
ns
tilingual synthetic datasets can enhance their were the same for high-quality voice conversion.
generalizability to out-of-domain languages. IndicSynth Metadata: IndicSynth contains sep-
arate metadata files for voice cloning done through
IndicSynth Generation: To generate synthetic each model for each target language3 . For the TTS
data for the mimicry subset, we fine-tuned the models, the metadata includes the target speaker
XTTS-v2 model on IndicSuperb for each of the fol- ID, the ID of the bonafide target voice sample, the
lowing 10 target languages: Bengali, Gujarati, Kan- target speaker’s gender, the transcript, and the ID
nada, Malayalam, Marathi, Odia, Punjabi, Tamil, of the generated synthetic audio clip. Similarly, for
Telugu, and Urdu2 . Fine-tuning XTTS-v2 and the VC models, the metadata includes the source
generating synthetic data using the same bonafide speaker ID, bonafide IndicSuperb source audio clip
dataset (IndicSuperb) helps ensure a high similarity ID, target speaker ID, bonafide IndicSuperb target
between synthetic and target voices. In contrast, audio clip ID, the gender of the speakers, and the ID
the diversity subset includes synthetic audios di- of the synthetic audio clip. Such metadata is also
rectly generated from TTS and VC models without valuable for studying gender bias in multilingual au-
fine-tuning on IndicSuperb. We utilized the pub- dio deepfake detection (ADD) (Xu et al., 2024; Ju
licly available Coqui VITS model (a TTS model et al., 2024). Overall, utilizing the bonafide Indic-
trained on undisclosed Bengali data) to generate Superb with the synthetic IndicSynth can facilitate
Bengali synthetic data for diversity subset (Eren multilingual ADD and anti-spoofing research.
and The Coqui TTS Team, 2021). Additionally, we
generated synthetic data for each of the 12 target 4 Evaluation of IndicSynth
languages using the publicly available XTTS-v2
4.1 IndicSynth for Audio DeepFake Detection
text-to-speech model and the FreeVC24 voice con-
version model (Eren and The Coqui TTS Team, Audio deepfake detection (ADD) models accept
2021), as illustrated in Table 1. The bonafide audios a speech recording as input and determine if it is
and transcripts were randomly chosen for Indic- bonafide (real) or synthetic (fake). These models
Synth creation. Also, we ensured that the genders are vital to prevent the spread of fake news or misin-
of the randomly chosen source and target speakers formation from synthetic audio. Thus, we urgently
need ADD models that are inclusive and general-
2
Code for fine-tuning XTTS-v2: [Link]
3
anhnh2002/XTTSv2-Finetuning-for-New-Languages IndicSynth’s directory structure is in Appendix (A).
22040
izable to unseen languages. A crucial first step in EER (%)
Language G. Model Aasist Aasist-L RawNet-2
developing generalizable ADD models is bench- XTTS-v2 70.125 56.150 56.737
marking state-of-the-art (SOTA) models on unseen Bengali FreeVC24 87.963 86.563 53.537
language datasets. Benchmarking can help inves- VITS 93.200 89.363 48.287
XTTS-v2 65.050 55.150 50.113
tigate potential biases in these models. Therefore, Gujarati
FreeVC24 86.163 86.888 53.425
we used IndicSynth to benchmark three state-of- Hindi
XTTS-v2 42.013 45.438 14.525
the-art (SOTA) publicly available ADD models: FreeVC24 81.775 81.913 48.513
XTTS-v2 55.563 49.188 42.425
Aasist, Aasist-L, and RawNet2 (Jung et al., 2022; Kannada
FreeVC24 73.30 78.950 50.874
Tak et al., 2021)4 . The benchmark ADD models are Malayalam
XTTS-v2 67.575 55.013 46.888
trained on an English dataset (LA partition of the FreeVC24 85.825 83.600 55.675
XTTS-v2 56.712 52.825 48.037
ASVspoof 2019 challenge) (Wang et al., 2020). We Marathi
FreeVC24 79.512 81.587 52.525
evaluated them on IndicSynth without fine-tuning. Odia
XTTS-v2 57.488 51.575 48.487
Setup: We created separate test sets for each FreeVC24 78.888 82.975 44.350
XTTS-v2 57.575 52.863 47.225
target language and generative model, as illustrated Punjabi
FreeVC24 81.925 82.225 53.775
in Table 2. Each set includes randomly chosen Sanskrit
XTTS-v2 33.438 38.775 7.95
4000 bonafide female voice samples, 4000 bonafide FreeVC24 84.238 86.563 58.35
XTTS-v2 61.725 51.188 51.650
male voice samples, 4000 synthetic female voice Tamil
FreeVC24 81.700 83.138 54.999
samples, and 4000 synthetic male voice samples. Telugu
XTTS-v2 54.275 52.700 46.275
The bonafide data was taken from IndicSuperb, FreeVC24 75.650 79.000 52.787
XTTS-v2 62.763 55.363 49.250
whereas synthetic data was taken from IndicSynth. Urdu
FreeVC24 78.088 79.825 50.438
We refer to them as IndicSynth-IndicSuperb sets. Table 2: Benchmarking audio deepfake detection (ADD)
Evaluation Metric: False Match Rate (FMR) models in IndicSynth-IndicSuperb test sets without do-
and False Non-Match Rate (FNMR) are widely main adaptation. For a given target language and a
used metrics for evaluating biometric systems. particular ADD model, the highest Equal Error Rate
FMR is the rate at which an ADD model incorrectly (EER%) achieved across various generative models is
highlighted in bold. When evaluated without domain
classifies synthetic audios as bonafide. In contrast,
adaptation, the benchmark ADD models achieve ele-
FNMR is the rate at which an ADD model incor- vated EER% on IndicSynth-IndicSuperb test sets. Train-
rectly classifies bonafide audios as synthetic. The ing ADD models on multilingual synthetic datasets,
FMR and FNMR values of an ADD model vary such as IndicSynth, can enhance their generalizability.
with classification thresholds. At a particular classi-
fication threshold, FMR becomes equal to FNMR.
The value of FMR when it becomes equal to the
FNMR is known as the Equal Error Rate (EER).
Equal Error Rate (EER) is a standard evaluation
metric for audio deepfake detection systems (Wang
et al., 2020; Liu et al., 2023). Therefore, we bench-
marked ADD models using EER(%). Lower EER
indicates that the models accurately distinguished
between bonafide and synthetic audios.
Figure 3: Receiver Operating Characteristic (ROC)
Observations: The Aasist and Aasist-L achieve
Curve for Malayalam IndicSynth-IndicSuperb test set
an EER of 0.83% and 0.99% on the LA evaluation created using XTTS-v2. Low Area Under the Curve
set of the ASVspoof 2019 (Jung et al., 2022). Sim- (AUC%) indicates poor discriminative power of ADD
ilarly, the RawNet-2 achieved an EER of 22.38% models.
in the DF track of the ASVspoof 2021 (Liu et al.,
2023). However, as illustrated in Table 2, these synthetic clips. Furthermore, we plotted the Re-
benchmark models achieved significantly higher ceiver Operating Characteristic (ROC) curves as il-
EERs on IndicSynth-IndicSuperb test sets. Ele- lustrated in Figure 3. The ROC curve demonstrates
vated EERs on these test sets indicate that the mod- that the ADD models achieve extremely low Area
els struggled to distinguish between bonafide and Under the Curve (AUC) scores on the Malayalam
4
test set created using the synthetic clips obtained
Aasist: [Link]
RawNet2: [Link] from the XTTS-v2 model. Low AUC indicates
2021/tree/main/DF/Baseline-RawNet2 poor discriminative power of the models. These ob-
22041
servations demonstrate a significant performance Language Source Accuracy ∆ Accuracy (%)
Bonafide 89.925 -
degradation of benchmark ADD models on unseen XTTS-v2 89.763 −0.162
Bengali
language test sets5 . Training such models on large- FreeVC24 90.338 +0.413
scale multilingual synthetic speech datasets can VITS 98.425 +8.500
Bonafide 98.612 -
potentially enhance their generalizability (Müller Gujarati XTTS-v2 96.762 −1.850
et al., 2024b). Therefore, IndicSynth is a valuable FreeVC24 96.475 −2.137
contribution towards generalizable ADD. Bonafide 92.250 -
Hindi XTTS-v2 86.175 −6.075
FreeVC24 85.525 −6.725
4.2 Linguistic Authenticity of IndicSynth Bonafide 88.550 -
Next, we investigate whether IndicSynth’s syn- Kannada XTTS-v2 84.800 −3.750
FreeVC24 85.638 −2.912
thetic speech recordings accurately capture the lin- Bonafide 97.425 -
guistic traits of the target languages. For this ex- Malayalam XTTS-v2 96.200 −1.225
FreeVC24 96.362 −1.063
periment, we created bonafide (IndicSuperb) and
Bonafide 94.900 -
synthetic (IndicSynth) test sets for each genera- Marathi XTTS-v2 89.725 −5.175
tive model and target language, as illustrated in FreeVC24 89.950 −4.950
Bonafide 78.388 -
Table 3. Each test set contains 8,000 audio clips. Punjabi XTTS-v2 66.600 −11.788
These sets contain an equal number of male and FreeVC24 65.938 −12.450
female voice samples. Subsequently, we evaluated Bonafide 41.350 -
Sanskrit XTTS-v2 9.050 −32.300
IndicSynth’s linguistic authenticity by running the FreeVC24 9.175 −32.175
publicly available VoxLingua107 ECAPA-TDNN Bonafide 97.500 -
spoken language identification model (Valk and Tamil XTTS-v2 94.500 −3.000
FreeVC24 94.812 −2.688
Alumäe, 2021; Ravanelli et al., 2021) on these sets6 . Bonafide 98.625 -
The model is trained on the VoxLingua107 dataset, Telugu XTTS-v2 96.100 −2.525
which includes speech recordings from 107 lan- FreeVC24 95.638 −2.987
Bonafide 39.900 -
guages (Valk and Alumae, 2021). Urdu XTTS-v2 33.975 −5.925
Observations: Table 3 illustrates the accuracy FreeVC24 33.763 −6.137
achieved on running the language identification
model through the test sets. We observe an ac- Table 3: Language identification results. We evaluated
IndicSynth’s linguistic authenticity by running language
curacy of more than 80% for most sets. Addi-
identification model through various test sets for each
tionally, we compared the accuracy difference be- generative model and target language (except Odia). We
tween bonafide and synthetic test sets defined as: observe above 80% accuracy in most test sets.
∆ Accuracy%=Accuracysynthetic −Accuracybonafide .
For most languages, the accuracy drop is below ECAPA-TDNN spoken language identification
10%. Interestingly, the accuracies of Bengali syn- model does not support Odia. However, the train-
thetic audios from FreeVC24 and VITS are higher ing set of the model covers 107 languages. There-
than bonafide audios, which indicates that these fore, its embeddings should capture linguistic traits
models are trained on diverse Bengali datasets. effectively. Thus, we obtained the 256-dimensional
language identification embeddings of the Odia
test sets and visualized them through t-SNE using
a perplexity of 40 (Munir et al., 2024), as shown
in Figure 4. Figure 4 shows no clear separation be-
tween the bonafide (IndicSuperb) and the synthetic
(IndicSynth) embeddings. The plot indicates that
IndicSynth’s Odia subset has effectively captured
Figure 4: t-SNE visualization of bonafide (IndicSuperb) linguistic traits of Odia7 .
and synthetic (IndicSynth) Odia dataset. The plot in-
dicates that the IndicSynth-Odia subset has effectively 4.3 Utility of the Mimicry Subset
captured the linguistic traits of Odia. Speaker verification systems accept two speech
Qualitative evaluation: The VoxLingua107 recordings as input and determine if they are from
5
the same speaker. The input speech recordings
Appendix (B) includes additional plots.
6 7
Language identification model: [Link] The t-SNE plots for Punjabi, Sanskrit, and Urdu are in
co/speechbrain/lang-id-voxlingua107-ecapa Appendix (D)
22042
Language SV Model Female Test Set Male Test Set Combined Test Set
ECAPA-TDNN 31.580 29.180 31.110
Bengali ResNet TDNN 23.960 22.960 25.160
X-Vector 43.860 44.520 43.639
ECAPA-TDNN 23.280 34.440 28.880
Gujarati ResNet TDNN 18.520 30.560 24.730
X-Vector 31.040 32.640 32.200
ECAPA-TDNN 25.880 22.700 24.360
Kannada ResNet TDNN 22.740 20.060 21.470
X-Vector 35.020 28.740 31.770
ECAPA-TDNN 30.140 33.680 32.070
Malayalam ResNet TDNN 30.700 31.480 31.170
X-Vector 38.460 38.240 38.520
ECAPA-TDNN 20.620 28.820 24.710
Marathi ResNet TDNN 17.680 25.140 21.600
X-Vector 28.260 31.160 31.080
ECAPA-TDNN 29.300 36.800 32.940
Odia ResNet TDNN 23.580 21.420 25.680
X-Vector 33.440 48.100 42.450
ECAPA-TDNN 25.800 33.420 29.610
Punjabi ResNet TDNN 23.360 30.640 27.050
X-Vector 32.330 36.160 34.350
ECAPA-TDNN 20.160 27.760 23.840
Tamil ResNet TDNN 17.820 25.500 21.540
X-Vector 28.280 31.860 30.070
ECAPA-TDNN 26.380 27.500 26.930
Telugu ResNet TDNN 23.320 25.140 24.550
X-Vector 31.760 32.120 32.180
Table 4: Investigating the vulnerability of state-of-the-art (SOTA) speaker verification models (SV) against imper-
sonation attacks. We observe elevated equal error rates (EER%) when the negative trial pairs contain IndicSynth’s
mimicry subset’s synthetic speech recordings and the target speaker’s bonafide speech from IndicSuperb. It suggests
that the mimicry subset audios closely mimic the bonafide target voices. Therefore, the mimicry subset of IndicSynth
is a valuable resource for enhancing the robustness of SOTA SV models.
form a trial pair. Such systems are vital in forensics, Methodology: We created speaker verification
business, e-commerce, and access control mecha- test sets, as illustrated in Table 4. Each set contains
nisms. However, the malicious use of voice cloning randomly generated 20,000 trial pairs with equal
models may lead to the generation of synthetic positives and negatives. A positive trial pair con-
speech recordings that closely mimic the target tains two bonafide (IndicSuperb) speech recordings
voice. Such synthetic recordings (audio spoofs) of the same target speaker, X. In contrast, a neg-
may be misused to deceive speaker verification ative trial pair contains a bonafide (IndicSuperb)
systems, leading to impersonation attacks against speech recording of a target speaker X and a syn-
the target speaker. Fine-tuning speaker verification thetic (IndicSynth) speech recording of X. Since
models on multilingual synthetic speech datasets each set contains an equal number of male and
can enhance their generalizability and robustness to female speaker trial pairs, we refer to them as com-
out-of-domain audio spoofs. Therefore, this exper- bined test sets. Each combined test set contains
iment explores the utility of IndicSynth’s mimicry 5000 bonafide female trial pairs, 5000 bonafide
subset for enhancing the robustness of speaker veri- male trial pairs, 5000 synthetic female trial pairs,
fication models. We evaluate whether the synthetic and 5000 synthetic male trial pairs. Subsequently,
audios of mimicry subset can deceive three pub- to evaluate gender bias in speaker verification mod-
licly available state-of-the-art (SOTA) speaker ver- els with respect to impersonation attacks, we also
ification models: ECAPA-TDNN, X-Vector, and split the combined test set and created separate
ResNet TDNN (Desplanques et al., 2020; Snyder male and female speaker test sets.
et al., 2018; Villalba et al., 2020)8 . Evaluation Metric: We evaluate the mimicry
8
ECAPA-TDNN:[Link] subset using EER. A higher EER indicates that the
speechbrain/spkrec-ecapa-voxceleb speaker verification model struggled to distinguish
X-Vector:[Link] between positive and negative trial pairs. It im-
spkrec-xvect-voxceleb
ResNet TDNN: [Link] plies that the synthetic speech recordings closely
spkrec-resnet-voxceleb mimic the target speaker’s bonafide voice sample
22043
in a negative trial pair. the t-SNE plots for Odia female and male voices.
Observations: The SOTA speaker verification For each plot, we randomly sampled 500 bonafide
models typically achieve an EER of less than and 500 synthetic clips of the same target speak-
10% when evaluated on unseen language test sets ers. Next, we created t-SNE plots with a perplexity
(Akram et al., 2024; Xia et al., 2019; Mandalapu of 40 using 80-dimensional Mel-Frequency Cep-
et al., 2021). However, as illustrated in Table 4, stral Coefficients (MFCC) features of these audios
we observed significantly elevated EERs ranging (Munir et al., 2024). MFCCs are biologically in-
from 21.470% to 43.639% on the combined test spired speech features that mimic the human au-
sets. Elevated EERs suggest that the speech record- ditory system. The proximity of the bonafide and
ings in IndicSynth’s mimicry subset closely mimic synthetic embeddings in t-SNE indicates that the
the bonafide (IndicSuperb) target voices. Further- mimicry subset’s synthetic audios closely mimic
more, we compared the EERs of the male and fe- the bonafide target voices9 .
male speaker test sets. The EERs of male and fe-
male speaker test sets for the Bengali, Malayalam, 5 Discussion
and Telugu test sets are comparable. However,
This section reflects on our rationale behind cre-
the EER values for Kannada female test sets are
ating mimicry and diversity subsets in IndicSynth.
higher than the male test sets (with absolute differ-
Additionally, we briefly review the potential utility
ences of 2.68% to 6.28% across the speaker veri-
of these subsets with our experimental results.
fication models). This observation indicates that
Rationale behind Mimicry Subset: Indic-
the Kannada female voices mimic the target speak-
Synth’s mimicry subset consists of synthetic voices
ers more closely than the Kannada male voices in
closely mimicking the target speaker’s bonafide
IndicSynth. Similarly, the EERs of the male test
voice. Such synthetic audios—also called au-
sets are higher than the female test sets for Gujarati,
dio spoofs—can deceive speaker verification sys-
Marathi, Odia, Punjabi, and Tamil. This observa-
tems, leading to impersonation attacks. Section
tion indicates that in IndicSynth, male voices in
4.3 demonstrates that speaker verification systems
these languages closely mimic the target speakers
are vulnerable to multilingual audio spoofs (Indic-
compared to female voices.
Synth’s mimicry subset). This observation under-
scores the potential utility of the mimicry subset
for developing multilingual anti-spoofing technolo-
gies.
Need for a Diversity Subset: Synthetic speech
is not only misused for impersonation but also for
spreading misinformation. Synthetic voices cir-
culating in social media to spread misinformation
Figure 5: The t-SNE plot of bonafide (IndicSuperb)
often do not mimic a specific target speaker’s voice.
Odia and IndicSynth’s mimicry subset’s female speak-
ers. The plot reveals the proximity of the bonafide and Instead, misinformation campaigns often involve
synthetic audios. diverse synthetic voices. Therefore, training audio
deepfake detection (ADD) models on a broad range
of synthetic voices is crucial for enhancing the ro-
bustness of these models. However, as we know,
the mimicry subset has a limited speaker diver-
sity as it only includes voices mimicking IndicSu-
perb speakers. Therefore, we introduced the diver-
sity subset in IndicSynth to incorporate a broader
Figure 6: The t-SNE plot of bonafide (IndicSuperb) range of synthetic voices beyond speaker mimicry
Odia and IndicSynth’s mimicry subset’s male speakers. into our dataset. Table 1 provides an overview of
The plot reveals the proximity of the bonafide and syn-
mimicry and diversity subsets.
thetic audios.
Utility of Diversity Subset: As described in Sec-
Qualitative Evaluation: For an extensive evalu-
tion 4.1, we evaluated three state-of-the-art (SOTA)
ation, we visualized the proximity of the bonafide
audio deepfake detection (ADD) models on Indic-
(IndicSuperb) and IndicSynth’s mimicry subset
through t-SNE. Figure 5 and Figure 6 represent 9
Additional t-SNE plots are in Appendix (C).
22044
Synth (both diversity and mimicry subsets) without deepfake detection research.
domain adaptation. These ADD models achieve This work opens up several avenues for research
low Equal Error Rates (EERs) on ASVSpoof chal- on linguistic biases in audio deepfake detection and
lenge datasets. However, we observed a significant anti-spoofing. For instance, IndicSynth can be used
performance degradation of these models on Indic- to investigate the underexplored problem of gender
Synth test sets, as illustrated by elevated EERs and biases in multilingual audio deepfake detection and
low Area Under the Curve (AUC) in Table 2 and for defense against impersonation attacks through
Figure 3. Munir et al. (2024) also reported a similar multilingual spoofs. The dataset is licensed under
observation. In their work, the authors proposed the CC BY-NC 4.0 license.
an Urdu audio deepfake detection dataset. Such
observations indicate a lack of generalizability of 7 Limitations
existing audio deepfake detection models to out-of-
domain languages. It suggests that training or fine- This work introduces IndicSynth, a novel, large-
tuning audio deepfake detection models on multi- scale multilingual synthetic speech dataset to fa-
lingual ADD datasets can potentially enhance the cilitate multlingual audio deepfake detection and
generalizability of these models to out-of-domain anti-spoofing research. We acknowledge that our
languages. Thus, bonafide IndicSuperb data com- work has the following limitations:
bined with the synthetic IndicSynth data (diversity IndicSynth’s Scope and Experimentation: In-
and mimicry subsets) can potentially serve as a dicSynth contains synthetic speech recordings for
valuable dataset for multilingual audio deepfake only 12 languages. Also, the mimicry subsets for
detection. Hindi and Sanskrit are absent in IndicSynth. In
the future, the dataset can be extended by cover-
6 Conclusions and Future Work ing more low-resourced languages and more voice
cloning models for dataset creation. Additionally,
This paper introduces IndicSynth, a novel large- we evaluated IndicSynth using sample test sets. We
scale multilingual synthetic speech dataset to fa- believe that the results from these sets indicate the
cilitate multilingual audio deepfake detection and overall dataset quality.
anti-spoofing research. The dataset contains about Absence of User Study: Ideally, the naturalness
4,000 hours of synthetic audio from 989 target of synthetic speech datasets should be evaluated
speakers, including 456 females and 533 males through a user study. The user study participants
for 12 low-resourced Indian languages. Addition- must be proficient in the target languages for au-
ally, IndicSynth includes rich metadata covering thentic results. However, for large-scale multilin-
the identifiers and gender information of the target gual datasets, such as IndicSynth, recruiting partic-
speakers. Thus facilitating research on gender bi- ipants proficient in low-resource languages is chal-
ases in audio deepfake detection and anti-spoofing. lenging. Thus, meeting our paper’s objectives, we
The dataset consists of mimicry and diversity sub- experimentally evaluated IndicSynth using state-
sets. The mimicry subset includes synthetic au- of-the-art speaker verification models, audio deep-
dios that closely mimic bonafide target voices. In fake detection models, and a language identifica-
contrast, the diversity subset contains a diverse tion model. Additionally, we have included t-SNE
set of realistic synthetic voices. Experimental re- plots in the appendix for a qualitative evaluation of
sults demonstrate that the synthetic audios of the our dataset.
mimicry subset can deceive state-of-the-art (SOTA) We highlight that the challenge of recruiting par-
speaker verification models through impersonation ticipants proficient in low-resourced languages is
attacks. Similarly, empirical results demonstrate not unique to IndicSynth. Instead, it is a common
poor performance of SOTA audio deepfake detec- issue faced by researchers working towards gener-
tion models on Indian language test sets. Further- ating multilingual datasets for social good. How-
more, qualitative and quantitative evaluation using ever, with around 7,000 global languages and rising
a SOTA language identification model validated cases of deepfake-related fraud, there is an urgent
the linguistic authenticity of our dataset. It turns need for multilingual synthetic datasets to facilitate
out that IndicSynth is a valuable contribution to research on multilingual audio deepfake detection.
preventing impersonation attacks on speaker verifi- Therefore, the absence of user studies should not
cation systems and facilitating multilingual audio hinder the release of such datasets. Instead, the
22045
community members with access to computational Jordan J. Bird and Ahmad Lotfi. 2023. Real-time de-
resources can contribute by constructing and releas- tection of ai-generated speech for deepfake voice
conversion. ArXiv, abs/2308.12734.
ing more multilingual datasets. Subsequently, the
members who can connect to native speakers of Brecht Desplanques, Jenthe Thienpondt, and Kris De-
those languages can contribute by conducting user muynck. 2020. ECAPA-TDNN: emphasized chan-
studies to evaluate human perception of synthetic nel attention, propagation and aggregation in TDNN
based speaker verification. In Interspeech 2020,
speech. pages 3830–3834. ISCA.
Despite these limitations, the dearth of multilin-
gual synthetic speech datasets makes IndicSynth a Gölge Eren and The Coqui TTS Team. 2021. Coqui
valuable resource that can facilitate research on TTS.
multilingual audio deepfake detection and anti- Joel Cameron Frank and Lea Schönherr. 2021. Wave-
spoofing. fake: A data set to facilitate audio deepfake detection.
ArXiv, abs/2111.02813.
8 Ethical Considerations Kurtis Haut, Caleb Wohn, Victor Antony, Aidan
Goldfarb, Melissa Welsh, Dillanie Sumanthiran,
Synthetic speech datasets are essential to advance M. Rafayet Ali, and Ehsan Hoque. 2022. Demo-
audio deepfake detection and anti-spoofing re- graphic feature isolation for bias research using deep-
search. However, we realize that such datasets can fakes. In Proceedings of the 30th ACM Interna-
inadvertently contribute to the refinement of au- tional Conference on Multimedia, MM ’22, page
6890–6897, New York, NY, USA. Association for
dio deepfake generation technologies by malicious Computing Machinery.
users. Therefore, responsible management of these
resources is crucial. Thus, we release IndicSynth Tahir Javed, Kaushal Bhogale, Abhigyan Raman,
under CC BY-NC 4.0, restricting commercial use Pratyush Kumar, Anoop Kunchukuttan, and Mitesh
Khapra. 2023. Indicsuperb: A speech processing uni-
of our dataset. Furthermore, IndicSynth is a syn- versal performance benchmark for indian languages.
thetic speech dataset generated from the publicly Proceedings of the AAAI Conference on Artificial
available bonafide IndicSuperb dataset. IndicSu- Intelligence, 37:12942–12950.
perb is licensed under the Creative Commons CC0
Yan Ju, Shu Hu, Shan Jia, George H. Chen, and Si-
license (“no rights reserved”). The CC0 license wei Lyu. 2024. Improving Fairness in Deepfake
allows users to freely build upon, reuse, or enhance Detection . In 2024 IEEE/CVF Winter Conference
the dataset without restriction. We strongly encour- on Applications of Computer Vision (WACV), pages
age the community to use IndicSynth for social 4643–4653, Los Alamitos, CA, USA. IEEE Com-
puter Society.
good and advance research on multilingual audio
deepfake detection and anti-spoofing. Jee-weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-jin
Shim, Joon Son Chung, Bong-Jin Lee, Ha-Jin Yu, and
Acknowledgments Nicholas Evans. 2022. Aasist: Audio anti-spoofing
using integrated spectro-temporal graph attention net-
This work is supported by the Infosys Centre for works. In ICASSP 2022 - 2022 IEEE International
Conference on Acoustics, Speech and Signal Process-
Artificial Intelligence (CAI) at IIIT-Delhi. We also ing (ICASSP), pages 6367–6371.
thank the SBILab at IIIT-Delhi for their helpful
discussions and support. Piotr Kawa, Marcin Plata, and Piotr Syga. 2022. Attack
agnostic dataset: Towards generalization and stabi-
lization of audio deepfake detection. pages 4023–
4027.
References
Hasam Khalid, Shahroz Tariq, and Simon S. Woo. 2021.
Ali Akram, Marija Stanojevic, Malikeh Ehghaghi, Fakeavceleb: A novel audio-video multimodal deep-
and Jekaterina Novikova. 2024. Zero-shot multi- fake dataset. ArXiv, abs/2108.05080.
lingual speaker verification in clinical trials. ArXiv,
abs/2404.01981. Pavel Korshunov and Sébastien Marcel. 2022. Im-
proving generalization of deepfake detection with
Zhongjie Ba, Qing Wen, Peng Cheng, Yuwei Wang, data farming and few-shot learning. IEEE Transac-
Feng Lin, Li Lu, and Zhenguang Liu. 2023. Trans- tions on Biometrics, Behavior, and Identity Science,
ferring audio deepfake detection capability across 4(3):386–397.
languages. In Proceedings of the ACM Web Confer-
ence 2023, WWW ’23, page 2033–2044, New York, Xuechen Liu, Xin Wang, Md Sahidullah, Jose Patino,
NY, USA. Association for Computing Machinery. Héctor Delgado, Tomi Kinnunen, Massimiliano
22046
Todisco, Junichi Yamagishi, Nicholas Evans, An- In Findings of the Association for Computational Lin-
dreas Nautsch, and Kong Aik Lee. 2023. Asvspoof guistics: NAACL 2024, pages 379–394, Mexico City,
2021: Towards spoofed and deepfake speech detec- Mexico. Association for Computational Linguistics.
tion in the wild. IEEE/ACM Transactions on Audio,
Speech, and Language Processing, 31:2507–2522. Divya Sharma and Arun Balaji Buduru. 2022. FAt-
Net: Cost-effective approach towards mitigating the
Hareesh Mandalapu, Thomas Møller Elbo, Raghavendra linguistic bias in speaker verification systems. In
Ramachandra, and Christoph Busch. 2021. Cross- Findings of the Association for Computational Lin-
lingual speaker verification: Evaluation on x-vector guistics: NAACL 2022, pages 1247–1258, Seattle,
method. In Intelligent Technologies and Applica- United States. Association for Computational Lin-
tions, pages 215–226, Cham. Springer International guistics.
Publishing.
David Snyder, Daniel Garcia-Romero, Alan McCree,
Sheza Munir, Wassay Sajjad, Mukeet Raza, Emaan Ab- Gregory Sell, Daniel Povey, and Sanjeev Khudanpur.
bas, Abdul Hameed Azeemi, Ihsan Ayyub Qazi, and 2018. Spoken language recognition using x-vectors.
Agha Ali Raza. 2024. Deepfake defense: Construct- In The Speaker and Language Recognition Workshop
ing and evaluating a specialized Urdu deepfake audio (Odyssey 2018), pages 105–111.
dataset. In Findings of the Association for Computa-
tional Linguistics: ACL 2024, pages 14470–14480, Sidharth T and Guhan T. 2024. Deepfake technology in
Bangkok, Thailand. Association for Computational social media: Social and legal implications in india.
Linguistics. IJFMR, 6(6).
Nicolas Müller, Nicholas Evans, Hemlata Tak, Philip Hemlata Tak, Jose Patino, Massimiliano Todisco, An-
Sperl, and Konstantin Böttinger. 2024a. Harder or dreas Nautsch, Nicholas Evans, and Anthony Larcher.
different? understanding generalization of audio 2021. End-to-end anti-spoofing with rawnet2. In
deepfake detection. pages 2705–2709. ICASSP 2021 - 2021 IEEE International Confer-
ence on Acoustics, Speech and Signal Processing
Nicolas M. Müller, Piotr Kawa, Wei Herng Choong, (ICASSP), pages 6369–6373.
Edresson Casanova, Eren Gölge, Thorsten Müller,
Piotr Syga, Philip Sperl, and Konstantin Böttinger. Jörgen Valk and Tanel Alumäe. 2021. VoxLingua107:
2024b. Mlaad: The multi-language audio anti- a dataset for spoken language recognition. In Proc.
spoofing dataset. In 2024 International Joint Confer- IEEE SLT Workshop.
ence on Neural Networks (IJCNN), pages 1–7.
Jorgen Valk and Tanel Alumae. 2021. Voxlingua107:
Xiaoke Qi, Hao Gu, Jiangyan Yi, Jianhua Tao, Yong A dataset for spoken language recognition. pages
Ren, Jiayi He, and Siding Zeng. 2024. Madd: A 652–658.
multi-lingual multi-speaker audio deepfake detection
dataset. In 2024 IEEE 14th International Symposium Jesús Villalba, Nanxin Chen, David Snyder, Daniel
on Chinese Spoken Language Processing (ISCSLP), Garcia-Romero, Alan McCree, Gregory Sell,
pages 466–470. Jonas Borgstrom, Leibny Paola García-Perera,
Fred Richardson, Réda Dehak, Pedro A. Torres-
Mouna Rabhi, Spiridon Bakiras, and Roberto Di Pietro. Carrasquillo, and Najim Dehak. 2020. State-of-the-
2024. Audio-deepfake detection: Adversarial attacks art speaker recognition with neural network embed-
and countermeasures. Expert Systems with Applica- dings in nist sre18 and speakers in the wild evalua-
tions, 250:123941. tions. Computer Speech & Language, 60:101026.
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Xin Wang, Junichi Yamagishi, Massimiliano Todisco,
Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Héctor Delgado, Andreas Nautsch, Nicholas Evans,
Subakan, Nauman Dawalatabad, Abdelwahab Heba, Md Sahidullah, Ville Vestman, Tomi Kinnunen,
Jianyuan Zhong, Ju-Chieh Chou, Sung-Lin Yeh, Kong Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai
Szu-Wei Fu, Chien-Feng Liao, Elena Rastorgueva, Peng, Hsin-Te Hwang, Yu Tsao, Hsin-Min Wang,
François Grondin, William Aris, Hwidong Na, Yan Sébastien Le Maguer, Markus Becker, Fergus Hen-
Gao, Renato De Mori, and Yoshua Bengio. 2021. derson, Rob Clark, Yu Zhang, Quan Wang, Ye Jia,
SpeechBrain: A general-purpose speech toolkit. Kai Onuma, Koji Mushika, Takashi Kaneda, Yuan
Preprint, arXiv:2106.04624. ArXiv:2106.04624. Jiang, Li-Juan Liu, Yi-Chiao Wu, Wen-Chin Huang,
Tomoki Toda, Kou Tanaka, Hirokazu Kameoka, In-
ARTH JUHUL SHAH, Ravindrakumar M. Purohit, gmar Steiner, Driss Matrouf, Jean-François Bonas-
Dharmendra H. Vaghera, and Hemant Patil. 2024. tre, Avashna Govender, Srikanth Ronanki, Jing-Xuan
MLADDC: Multi-lingual audio deepfake detection Zhang, and Zhen-Hua Ling. 2020. Asvspoof 2019: A
corpus. In Audio Imagination: NeurIPS 2024 Work- large-scale public database of synthesized, converted
shop AI-Driven Speech, Music, and Sound Genera- and replayed speech. Computer Speech & Language,
tion. 64:101114.
Divya Sharma. 2024. EcoSpeak: Cost-efficient bias mit- Wei Xia, Jing Huang, and John H.L. Hansen. 2019.
igation for partially cross-lingual speaker verification. Cross-lingual text-independent speaker verification
22047
IndicSynth/
using unsupervised adversarial discriminative do- |-- <language>/
main adaptation. In ICASSP 2019 - 2019 IEEE Inter- | |-- XTTS_v2/
national Conference on Acoustics, Speech and Signal | | |-- Male/
Processing (ICASSP), pages 5816–5820. | | | |-- <speaker_id>/
| | | | |-- <clip_id>.wav
Yuankun Xie, Haonan Cheng, Yutian Wang, and Long | | | | `-- ...
Ye. 2024. Domain generalization via aggregation | | |-- Female/
and separation for audio deepfake detection. IEEE | | | |-- <speaker_id>/
Transactions on Information Forensics and Security, | | | | |-- <clip_id>.wav
19:344–358. | | | | `-- ...
| | `-- [Link]
Ying Xu, Philipp Terhörst, Marius Pedersen, and Kiran | |-- FreeVC24/
Raja. 2024. Analyzing fairness in deepfake detection | | |-- Male/
with massively annotated databases. IEEE Transac- | | | |-- <speaker_id>/
tions on Technology and Society, 5(1):93–106. | | | | |-- <clip_id>.wav
| | | | `-- ...
Jiangyan Yi, Ruibo Fu, Jianhua Tao, Shuai Nie, Haoxin | | |-- Female/
Ma, Chenglong Wang, Tao Wang, Zhengkun Tian, | | | |-- <speaker_id>/
Ye Bai, Cunhang Fan, Shan Liang, Shiming Wang, | | | | |-- <clip_id>.wav
| | | | `-- ...
Shuai Zhang, Xinrui Yan, Le Xu, Zhengqi Wen, and | | `-- [Link]
Haizhou Li. 2022. Add 2022: the first audio deep
synthesis detection challenge. In ICASSP 2022 -
2022 IEEE International Conference on Acoustics, Figure 7: IndicSynth directory structure.
Speech and Signal Processing (ICASSP), pages 9216–
9220.
B IndicSynth for Audio DeepFake
Jiangyan Yi, Jianhua Tao, Ruibo Fu, Xinrui Yan, Chen- Detection: Additional Plots
glong Wang, Tao Wang, Chu Yuan Zhang, Xiao-
hui Zhang, Yan Zhao, Yong Ren, Leling Xu, Jun Figures 8–56 illustrate the Receiver Operating
Zhou, Hao Gu, Zhengqi Wen, Shan Liang, Zheng
Lian, Shuai Nie, and Haizhou Li. 2023. Add 2023:
Characteristic (ROC) Curves and the Detection Er-
the second audio deepfake detection challenge. In ror Trade-off (DET) Curves for IndicSynth test sets
DADA@IJCAI. created using various generative models. The plots
indicate that the benchmark audio deepfake detec-
Jiangyan Yi, Chu Yuan Zhang, Jianhua Tao, Cheng-
long Wang, Xinrui Yan, Yong Ren, Hao Gu, and tion models lack generalizability on out-of-domain
Junzuo Zhou. 2024. Add 2023: Towards audio deep- (Indian) language test sets.
fake detection and analysis in the wild. Preprint,
arXiv:2408.04967.
Mohammed Yousif, Jonat John Mathew, Huzaifa Pal-
lan, Agamjeet Singh Padda, Syed Daniyal Shah, Sara
Adamski, Madhu Reddiboina, and Arjun Pankajak-
shan. 2024. Enhancing generalization in audio deep-
fake detection: A neural collapse based sampling and
training approach. Preprint, arXiv:2404.13008.
Yi Zhu, Surya Koppisetti, Trang Tran, and Gaurav
Bharaj. 2024. Slim: Style-linguistics mismatch
model for generalized audio deepfake detection.
ArXiv, abs/2407.18517.
Figure 8: The Receiver Operating Characteristic (ROC)
A IndicSynth Directory Structure Curve for the Bengali test set created using XTTS-v2.
Low Area Under the Curve (AUC) values of audio
Figure 7 shows the directory structure of Indic- deepfake detection models indicate poor discriminative
Synth. The dataset contains a folder for each target power of these models on our test set.
language. Each language folder includes subfold-
ers for generative models: XTTS_v2, FreeVC24,
or the VITS. Each generative model folder has a
metadata file and subfolders for anonymous target
speaker IDs. Within each target speaker’s folder,
there are synthetic audio clips. The speaker IDs
used in the IndicSynth are the same as those used
in IndicSuperb.
22048
Figure 9: The Detection Error Trade-off (DET) Curve Figure 12: The Receiver Operating Characteristic
for the Bengali test set created using XTTS-v2. The (ROC) Curve for the Bengali test set created using
trend of DET curves towards the upper right indicates VITS. Low Area Under the Curve (AUC) values of
the poor capability of the audio deepfake detection mod- audio deepfake detection models indicate poor discrimi-
els in distinguishing between bonafide and synthetic native power of these models on our test set.
audios.
Figure 10: The Receiver Operating Characteristic Figure 13: The Detection Error Trade-off (DET) Curve
(ROC) Curve for the Bengali test set created using for the Bengali test set created using VITS. The trend of
FreeVC24. Low Area Under the Curve (AUC) values DET curves towards the upper right indicates the poor
of audio deepfake detection models indicate poor dis- capability of the audio deepfake detection models in
criminative power of these models on our test set. distinguishing between bonafide and synthetic audios.
22049
Figure 15: The Detection Error Trade-off (DET) Curve
for the Gujarati test set created using XTTS-v2. The Figure 18: The Receiver Operating Characteristic
trend of DET curves towards the upper right indicates (ROC) Curve for the Hindi test set created using XTTS-
the poor capability of the audio deepfake detection mod- v2. Low Area Under the Curve (AUC) values of audio
els in distinguishing between bonafide and synthetic deepfake detection models indicate poor discriminative
audios. power of these models on our test set.
Figure 17: The Detection Error Trade-off (DET) Curve Figure 20: Receiver Operating Characteristic (ROC)
for the Gujarati test set created using FreeVC24. The Curve for the Hindi IndicSynth-IndicSuperb test set
trend of DET curves towards the upper right indicates created using FreeVC24. Low Area Under the Curve
the poor capability of the audio deepfake detection mod- (AUC%) indicates poor discriminative power of ADD
els in distinguishing between bonafide and synthetic models.
audios.
22050
Figure 21: The Detection Error Trade-off (DET) Curve Figure 24: The Receiver Operating Characteristic
for the Hindi test set created using FreeVC24. The trend (ROC) Curve for the Kannada test set created using
of DET curves towards the upper right indicates the FreeVC24. Low Area Under the Curve (AUC) values
poor capability of the audio deepfake detection models of audio deepfake detection models indicate poor dis-
in distinguishing between bonafide and synthetic audios. criminative power of these models on our test set.
Figure 23: The Detection Error Trade-off (DET) Curve Figure 26: The Detection Error Trade-off (DET) Curve
for the Kannada test set created using XTTS-v2. The for the Malayalam test set created using XTTS-v2. The
trend of DET curves towards the upper right indicates trend of DET curves towards the upper right indicates
the poor capability of the audio deepfake detection mod- the poor capability of the audio deepfake detection mod-
els in distinguishing between bonafide and synthetic els in distinguishing between bonafide and synthetic
audios. audios.
22051
Figure 27: The Receiver Operating Characteristic Figure 30: The Detection Error Trade-off (DET) Curve
(ROC) Curve for the Malayalam test set created using for the Marathi test set created using XTTS-v2. The
FreeVC24. Low Area Under the Curve (AUC) values trend of DET curves towards the upper right indicates
of audio deepfake detection models indicate poor dis- the poor capability of the audio deepfake detection mod-
criminative power of these models on our test set. els in distinguishing between bonafide and synthetic
audios.
22052
Figure 33: The Receiver Operating Characteristic Figure 36: The Detection Error Trade-off (DET) Curve
(ROC) Curve for the Odia test set created using XTTS- for the Odia test set created using FreeVC24. The trend
v2. Low Area Under the Curve (AUC) values of audio of DET curves towards the upper right indicates the
deepfake detection models indicate poor discriminative poor capability of the audio deepfake detection models
power of these models on our test set. in distinguishing between bonafide and synthetic audios.
22053
Figure 39: The Receiver Operating Characteristic Figure 42: The Detection Error Trade-off (DET) Curve
(ROC) Curve for the Punjabi test set created using for the Sanskrit test set created using XTTS-v2. The
FreeVC24. Low Area Under the Curve (AUC) values trend of DET curves towards the bottom left indicates
of audio deepfake detection models indicate poor dis- the capability of the audio deepfake detection models in
criminative power of these models on our test set. distinguishing between bonafide and synthetic audios.
22054
Figure 45: The Receiver Operating Characteristic
(ROC) Curve for the Tamil test set created using XTTS-
v2. Low Area Under the Curve (AUC) values of audio Figure 47: The Receiver Operating Characteristic
deepfake detection models indicate poor discriminative (ROC) Curve for the Tamil test set created using
power of these models on our test set. FreeVC24. Low Area Under the Curve (AUC) values
of audio deepfake detection models indicate poor dis-
criminative power of these models on our test set.
22055
Figure 50: The Detection Error Trade-off (DET) Curve Figure 53: The Receiver Operating Characteristic
for the Telugu test set created using XTTS-v2. The trend (ROC) Curve for the Urdu test set created using XTTS-
of DET curves towards the upper right indicates the poor v2. Low Area Under the Curve (AUC) values of audio
capability of the audio deepfake detection models in deepfake detection models indicate poor discriminative
distinguishing between bonafide and synthetic audios. power of these models on our test set.
22056
Figure 58: The t-SNE plot of bonafide (IndicSuperb)
Figure 56: The Detection Error Trade-off (DET) Curve Gujarati and IndicSynth’s mimicry subset’s male speak-
for the Urdu test set created using FreeVC24. The trend ers. The plot reveals the proximity of the bonafide and
of DET curves towards the upper right indicates the synthetic audios.
poor capability of the audio deepfake detection models
in distinguishing between bonafide and synthetic audios.
E Costs
This section highlights the costs of various experi-
ments conducted for this study in terms of carbon
emissions, electricity consumption, and execution
time. The data generation and experimentation
were done using the NVIDIA RTX A6000 GPU.
Below are the details:
Figure 73: The t-SNE plot of bonafide (IndicSuperb)
Urdu and IndicSynth’s mimicry subset’s male speakers. 1. Cost of fine-tuning XTTS-v2: Fine-tuning
The plot reveals the proximity of the bonafide and syn-
the XTTS-v2 for one epoch for a particular
thetic audios.
target language takes approximately 1.5 hours.
22059
In those 1.5 hours, the process causes approx- 22,150,912 parameters and occupies 364 MB
imately 0.403 kgCO2eq carbon emission and of memory.
consumes 0.564 kWh of electricity. This re- 8. X-Vector speaker verification model: It took
sult is for the training set containing 70,692 about 1.5 minutes to generate ResNet TDNN
audio clips. We fine-tuned each model for embeddings for 10,000 audio clips. The
45 epochs. The XTTS-v2 model contains model caused approximately 0.003 kgCO2eq
470,751,571 parameters. carbon emission and energy consumption of
2. XTTS-v2: It takes about 45 minutes to gen- 0.004 kWh for 10,000 embeddings. The X-
erate 1,000 audio clips from the XTTS-v2 Vector model has 8,172,473 parameters and
model for Hindi. The average duration of a occupies 300 MB of memory.
Hindi audio clip in IndicSuperb is 2.65 sec- 9. AASIST Audio Deepfake Detection Model:
onds. For generating 1000 synthetic audios us- The AASIST model has 297,866 parameters
ing the XTTS-v2 model, approximately 0.106 and occupies 264 MB of memory.
kgCO2eq carbon emission and 0.148 kWh 10. AASIST-L Audio Deepfake Detection
of electricity are consumed. The XTTS-v2 Model: The AASIST model has 85,306 pa-
model has 470,751,571 parameters and occu- rameters and occupies 262 MB of memory.
pies 2,110 MB of memory. 11. RawNet-2 Audio DeepFake Detection
3. FreeVC24: It takes about 4 minutes to gen- Model: The RawNet-2 model contains
erate 1000 audio clips from the FreeVC24 17,623,671 parameters and occupies 67.2 MB
model for Hindi. The average duration of of disk space.
a Hindi audio clip in IndicSuperb is 2.65
seconds. For generating 1000 synthetic au- F Tools and Software used
dios using the FreeVC24 model, approxi- We used the following tools and software for this
mately 0.010 kgCO2eq carbon emissions and study (other than the ones already cited in the pa-
0.014 kWh of electricity are consumed. The per):
FreeVC24 model contains 356,216,448 pa-
rameters and occupies 1690 MB of memory. 1. We used Grammarly and ChatGPT for better
4. VITS: The VITS model has 83,050,540 pa- sentence construction at occasional places and
rameters and occupies 586 MB of memory. to enhance clarity in our draft.
5. Language identification: It took about 2 min- 2. We used [Link] and the matplotlib for dia-
utes to run the VoxLingua107 ECAPA-TDNN grams.
spoken language identification model on a test 3. We used Librosa to generate the MFCC fea-
set containing 8,000 audio clips. The model tures.
caused approximately 0.004 kgCO2eq carbon 4. We used Pytorch for experimentation: version
emission and energy consumption of 0.006 2.5.1+cu124.
kWh.
6. ResNet TDNN speaker verification model: G Licenses
It took about 4 minutes to generate ResNet In this section, we specify the licenses of the
TDNN embeddings for 10,000 audio clips. datasets and models that we have used for this
The model caused approximately 0.012 study.
kgCO2eq carbon emission and energy con-
sumption of 0.018 kWh for 10,000 em- 1. IndicSuperb: Creative Commons CC0 license
beddings. The ResNet TDNN model has ("no rights reserved").
17,282,816 parameters and occupies 334 MB 2. Aasist, Aasist-L, and the RawNet-2: The
of memory. ADD models used are licensed under the MIT
7. Ecapa-TDNN speaker verification model: License.
It took about 3 minutes to generate ResNet 3. The speechbrain models (Ecapa-TDNN,
TDNN embeddings for 10,000 audio clips. ResNet TDNN, X-Vector, VoxLingua107
The model caused approximately 0.007 ECAPA-TDNN spoken language identifica-
kgCO2eq carbon emission and energy con- tion model) is licensed under the Apache Li-
sumption of 0.009 kWh for 10,000 em- cense 2.0.
beddings. The Ecapa-TDNN model has
22060