0% found this document useful (0 votes)
13 views18 pages

Forensic Speech Recognition in Child Abuse

This document proposes using two speech recognition systems, Wav2vec2.0 and Whisper, to analyze audio data from online sources for evidence of child abuse. It aims to develop an AI-powered platform to help law enforcement agencies process growing amounts of multimedia data in a timely manner. The systems are tested on various languages and scenarios. Federated learning strategies are also explored to maintain privacy for law enforcement agencies while improving system performance. Test results found word error rates between 11-25% and true positive rates for keyword spotting between 82-98%, depending on the language. Federated learning showed potential to maintain or improve performance compared to centralized models.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views18 pages

Forensic Speech Recognition in Child Abuse

This document proposes using two speech recognition systems, Wav2vec2.0 and Whisper, to analyze audio data from online sources for evidence of child abuse. It aims to develop an AI-powered platform to help law enforcement agencies process growing amounts of multimedia data in a timely manner. The systems are tested on various languages and scenarios. Federated learning strategies are also explored to maintain privacy for law enforcement agencies while improving system performance. Test results found word error rates between 11-25% and true positive rates for keyword spotting between 82-98%, depending on the language. Federated learning showed potential to maintain or improve performance compared to centralized models.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.

v1

Disclaimer/Publisher’s Note: The statements, opinions, and data contained in all publications are solely those of the individual author(s) and
contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting
from any ideas, methods, instructions, or products referred to in the content.

Article
Novel Speech Recognition Systems Applied to Forensics within
Child Exploitation: Wav2vec2.0 vs. Whisper
Juan Camilo Vasquez-Correa1 , Aitor Alvarez-Muniain1

1 Fundación Vicomtech, Basque Research and Technology Alliance (BRTA), Mikeletegi 57, 20009 Donostia-San
Sebastián (Spain)
* Correspondence: {jcvasquez,aalvarez}@[Link] (J.C.V., A.A.)

Abstract: The growth in online child exploitation material is a significant challenge for European 1

Law Enforcement Agencies (LEAs). One of the most important sources of such online information 2

corresponds to audio material that needs to be analyzed to find evidence in a timely and practical 3

manner. That is why LEAs require a next-generation AI-powered platform to process audio data 4

from online sources. We propose the use of speech recognition and keyword spotting to transcribe 5

audiovisual data and to detect the presence of keywords related to child abuse. The considered 6

models are based on two of the most accurate neural-based architectures to date: Wav2vec2.0 7

and Whisper. The systems are tested under an extensive set of scenarios in different languages. 8

Additionally, keeping in mind that obtaining data from LEAs is very sensitive, we explore the use of 9

federated learning to have more robust systems for the addressed application, while maintaining 10

the privacy of the data to LEAs. The considered models achieved a word error rate between 11% 11

and 25%, depending on the language. In addition, the systems are able to recognize a set of spotted 12

words with true positives rates between 82% and 98%, depending on the language. Finally, federated 13

learning strategies show that they can maintain and even improve the performance of the systems 14

when compared to centralized trained models. The proposed systems sit the basis for an AI-powered 15

platform for automatic analysis of audio in the context of forensic applications within child abuse. 16

The use of federated learning is also promising for the addressed scenario, where data privacy is an 17

important issue to be managed. 18

Keywords: Speech Recognition; Keyword Spotting; Child abuse; Federated Learning; Whisper; 19

Wav2vec2.0 20

1. Introduction 21

The growth in online child exploitation and abuse material is a significant challenge 22

for European Law Enforcement Agencies (LEAs). Currently, the revision of online material 23

about child abuse exceeds the capacity of LEAs to respond in a practical and timely manner. 24

One of the most important sources of information that needs to be analyzed to find evidence 25

about child abuse corresponds to audiovisual material from multimedia content. With the 26

aim to safeguard victims, prosecute offenders and limit the spread of online child abuse 27

related material, LEAs need a next-generation AI-powered platform to process multimedia 28

data from online sources. One of the main goals of the GRACE project1 is to develop robust 29

AI-based technology to equip LEAs with the aforementioned platform. Two of the core 30

applications to be incorporated correspond to automatic speech recognition (ASR) and 31

keyword spotting (KWS) in order to accurately transcribe audiovisual online material, and 32

to detect the presence of specific keywords about child abuse in the transcriptions. 33

Within this context, ASR technology has been applied in different forensic scenarios. 34

For instance, to collect evidence via the examination of electronic devices [1], or to analyze 35

1 [Link]

© 2022 by the author(s). Distributed under a Creative Commons CC BY license.


Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

2 of 18

multimedia content related to specific threats [2,3]. Nevertheless, the successful implemen- 36

tation of an ASR system in forensics introduces a series of issues to be solved, which are not 37

present in other domains where ASR is applied. For instance, it is common to find audio 38

coming from different sources, which are highly affected by background noise, overlapping 39

speakers, audio reverberation, among other factors. All these aspects affect the quality of 40

the obtained transcription and the capability of the system to detect specific keywords. 41

Although the aforementioned problems, recent advances in ASR have introduced 42

novel end-to-end architectures [4] that have shown to be accurate enough in those adverse 43

conditions. The core idea of end-to-end models is to directly map the input speech signal 44

to character sequences and therefore greatly simplify training, fine-tuning and inference [5– 45

9]. Two main approaches are distinguished in the literature to train end-to-end ASR 46

systems: fully supervised or self-supervised models. Regarding the first group, NVIDIA 47

proposed Quartznet [10] with the aim to build a competitive but lighter end-to-end ASR 48

model. The architecture consists of multiple blocks of 1D convolutions stacked with 49

residual connections. The model has been trained and tested on the Common Voice 50

corpus, achieving Word Error Rates (WERs) between 7.7% and 12.5%, depending on the 51

language [11]. A Quartznet model also produced WERs of 19.2% and 18.3% in French 52

and Spanish language multimedia data, respectively, from the MediaSpeech corpus [12]. 53

Researchers from NVIDIA recently proposed Citrinet [13] as an evolution of Quartznet. The 54

model consists of a residual network formed by 1D time-channel separable convolutions 55

combined with a sub-word encoding and a squeeze-and-excitation mechanism [14]. The 56

authors reported a WER of 5.6% on the TEDLIUMv2 corpus. Another architecture that 57

have proven to be accurate in many ASR benchmark scenarios is the Recurrent Neural 58

Network Transducer (RNN-T) [15]. The RNN-T is formed by three main blocks: (1) an 59

encoder network that receives input acoustic frames and produces high-level speech 60

representations, (2) a predictor that acts as a decoder by processing the previous produced 61

token, and (3) a joint network that combines the outputs from the two previous blocks and 62

produces the distribution of the next predicted token or blank symbol. Recent models based 63

on RNN-T achieved a WER 14.0% in the TEDLIUMv2 corpus [16]. 64

Contrary to fully supervised models, recent studies are focused on the use of big 65

acoustic models trained with self-supervised learning methods and a large amount of 66

unlabelled data. Researchers from Meta AI demonstrated the capabilities of this type of 67

models by introducing Wav2Vec2.0 [17]. This system outperformed many benchmark 68

results, especially when considering ASR for low-resource languages in the Common Voice 69

corpus [18]. Particularly, the authors in [19] considered a Wav2vec2.0 model combined with 70

their proposed language modeling approach, and achieve state-of-the-art results in the 71

German Common Voice corpus, with a WER of 3.7%. Wav2Vec2.0-based models have also 72

been successfully tested in more adverse acoustic environments such as in multimedia Por- 73

tuguese data from the CORAA database [20]. Due to these reasons, Wav2Vec2.0 has become 74

one of the most considered neural-based models for ASR. Self-supervised approaches like 75

Wav2Vec2.0 are challenging because there is not a predefined lexicon for the input sound 76

units during the pre-training phase. Moreover, sound units have variable length with no 77

explicit segmentation [21]. With the aim to solve such issues, Meta AI released HuBERT 78

as a new approach to learn self-supervised speech representations [22]. The combination 79

of Convolutional and Transfomer networks from Wav2Vec2.0 and HuBERT has achieved 80

state-of-the-art results in many ASR scenarios. With the aim to combine the best features 81

from both type of networks in a single neural block, researchers from Google introduced the 82

"convolutional augmented Transfomer" or Conformer [23]. A Conformer network achieved 83

a WER of 7.2% in the TEDLIUMv2 corpus [24]. 84

Self-supervised audio encoders like Wav2Vec2.0, HuBERT, and Conformers learn high 85

quality audio representations. However, due to its unsupervised pre-training nature, they 86

lack a proper decoding to transform such representations into usable outputs. This is why 87

a fine-tuning stage is always necessary in order to accurately implement models for ASR 88

or audio classification. With the aim to solve the aforementioned issue, researchers from 89
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

3 of 18

OpenAI recently proposed "Whisper" [25]. Whisper is a sequence-to-sequence Transformer 90

trained in a fully supervised manner, using up to 680,000 hours of labelled audio from 91

internet. The model has achieved state-of-the-art WER results in many benchmark datasets 92

for ASR, including librispeech, TEDLIUM, Common Voice, among others. 93

There are two main issues that appear when designing ASR solutions for forensic 94

scenarios: The first one is related to find the most appropriate neural architecture from the 95

ones previously described in order to deal with different acoustic environments. The second 96

one is related to data privacy and protection [26]. Generally, obtaining operative data from 97

LEAs for the addressed scenario is not possible. In this context, Federated Learning (FL) 98

has emerged as an alternative to train machine learning models over remote devices such 99

as mobile phones or remote data-centers in a non-centralized manner, preserving data 100

privacy [27–30]. The procedure is as follows: LEAs operative data are stored in on-premise 101

data servers. Then, FL strategies aim to transfer only local model updates to a central 102

server, keeping LEAs data private. The central server aggregates information obtained 103

from multiple clients i.e., LEAs, and updates a central model that is transmitted back to the 104

clients for their consumption. FL has been applied to train robust federated acoustic models 105

for ASR [31–33] and KWS [34]. In [32] the authors proposed a client adaptive federated 106

training to mitigate data heterogeneity when training ASR models. The proposed system 107

achieved a similar WER with respect to the obtained one using a fully centralized training. 108

In [33] the authors proposed a strategy to compensate Non Independent and Identically 109

Distributed (non-IID) data in federated training of ASR systems. The proposed strategy 110

involved random client data sampling, which resulted in a cost-quality trade-off. The 111

optimization of such a trade-off led to obtain ASRs with similar WERs than the obtained by 112

training centralized systems. The authors in [34] demonstrated the capabilities of federated 113

training to obtain robust KWS systems locally trained on edge devices like smartphones, 114

reaching similar accuracies when compared with centralized trained models. 115

According to the reviewed literature, the two main paradigms and solutions for ASR 116

to date include self-supervised models based on Wav2Vec2.0 and fully supervised models 117

such as Whisper. This work considered and compared these two approaches to test their 118

capabilities to perform robust ASR and KWS in a large set of test scenarios. We also 119

evaluated the use of FL in the context where different LEAs can share a common ASR and 120

KWS system keeping the privacy of their data. In summary, the main contributions of this 121

paper are four-folds: 122

1. We performed an extensive comparison between two of the most accurate neural- 123

based ASR architectures to date: a fine-tuned version of Wav2Vec2.0 and Whisper. 124

The evaluation is performed in many scenarios, but paying special attention to cor- 125

pora coming from multimedia content. The models are tested in data from seven 126

indo-European Languages, including English, Spanish, German, French, Italian, Por- 127

tuguese, and Polish. This evaluation can be useful as well in other domains besides 128

ASR forensics, making our contribution open and viable in other scenarios. 129

2. We created and released an in domain corpus that includes specific keywords of child 130

abuse domain, and a set of accompanying audios where the keywords are present. 131

The included audios are selected from open available corpora used in the literature. 132

The created corpus can be used as a benchmark to test ASRs in non-controlled acoustic 133

conditions. 134

3. The two neural architectures are compared as well in the created corpora within the 135

scope of child abuse forensics. To the best of our knowledge, this is the first study that 136

comprises the use of open ASR solutions and their capabilities to recognize specific 137

words within a forensic domain. 138

4. We validated the use of FL strategies to train ASR systems in the context of forensic 139

applications. The core idea is that different LEAs can share a common model but 140

keeping the privacy of their data. 141

The rest of the paper is distributed as follows. Section 2 details different technical 142

aspects of Wav2Vec2.0 and Whisper architectures for ASR. Section 3 describes the consid- 143
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

4 of 18

ered corpora to test the ASR systems, and the process to deliver an in domain corpus for 144

KWS in the context of forensics. Section 4 describes the pilot study on the use of FL for the 145

addressed application. Section 5 displays the main results obtained regarding ASR, KWS, 146

and FL. Section 6 discusses the main insights obtained from the results. Finally, Section 7 147

shows the main conclusion derived from this work. 148

2. Methods 149

We considered two of the most accurate neural-based ASR architectures to date: (1) 150

Wav2vec2.0, which is trained following a self-supervised paradigm, and (2) Whisper, which 151

is trained following a fully supervised strategy. Details about each model are found in the 152

following sub-sections. 153

2.1. Wav2vec2.0 154

Wav2vec2.0 [17] is a self-supervised end-to-end architecture based on convolutional 155

and Transformer layers (see Figure 1). The model encodes raw audio waveforms χ into 156

latent speech representations z1 , . . . , z T via a multi-layer convolutional feature encoder 157

f : χ → Z. These latent representations fed a Transformer-masked network g : Z → C. The 158

Transformer network initially quantise the continuous representations, forming a discrete 159

set of outputs q1 , . . . , q T that represent targets in the self-supervised learning objective [17, 160

35]. Those quantised representations are then contextualised using the attention blocks from 161

the Transformer module, obtaining a set of discrete contextual representations c1 , . . . , c T . 162

The feature encoder is formed by seven convolutional blocks with 512 channels, strides 163

of {5, 2, 2, 2, 2, 2, 2} and kernel widths of {10, 3, 3, 3, 3, 2, 2}. The Transformer network is 164

formed by 24 blocks, a model dimension of 1024, an inner dimension of 4096 and a total of 165

16 attention heads. 166

Figure 1. Wav2vec2.0 architecture representation. The raw audio signal is mapped to speech
representations that are fed into a Transformer network to output context representations. Figure
based on the one presented in [17].

We considered a pre-trained Wav2vec2.0 acoustic model based on the Wav2Vec2- 167

XLS-R-300M model, which is available via Hugginface2 . The model was pre-trained in 168

a self-supervised manner using 436k hours of unlabelled speech data in 128 languages 169

2 [Link]
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

5 of 18

from the VoxPopuli [36], Multilingual librispeech (MLS) [37], Common Voice [38], BABEL, 170

and VoxLingua107 [39] corpora. The Wav2Vec2-XLS-R-300M is one of the different versions 171

of the Meta AI’s XLS-R multilingual model [40] composed by 300 million parameters. The 172

multilingual pre-trained model was fine-tuned with labelled speech data (see Section 3.1) in 173

seven languages: English, German, French, Spanish, Italian, Portuguese, and Polish. Each 174

model was trained for 50 epochs, with a batch size of 2, 16 gradient accumulation steps, 175

and a learning rate of 5 × 10−5 , which is warmed up during the initial 10% of the training. 176

The trained acoustic representations are decoded using a Connectionist Temporal 177

Classification (CTC) layer with a beam-search decoding strategy (beam-width=256). The 178

CTC decoding include the use of separate 3-gram language models that are trained using 179

large text corpora, and which are included in the decoding with weights of α = 0.5 and 180

β = 1.5. 181

2.2. Whisper 182

Whisper is a recently introduced ASR system by OpenAI [25]. Contrary to Wav2vec2.0, 183

Whisper is trained in a fully supervised manner, using up to 680k hours of labelled speech 184

data from multiple sources. The model is based on an encoder-decoder Transformer, which 185

is fed by 80-channel log-Mel spectrograms. The encoder is formed by two convolution 186

layers with a kernel size of 3, followed by a sinusoidal positional encoding, and a stacked 187

set of Transformer blocks. The decoder uses the learned positional embeddings and the 188

same number of Transformer blocks from the encoder. Figure 2 illustrates the general 189

Whisper architecture. Different pre-trained models are available with variations in the 190

number of layers and attention heads. We considered the "Whisper-large" model, which 191

consists of 1550 million parameter distributed in 32 layers and 20 attention heads. The 192

model is available via Huggingface3 . 193

3 [Link]
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

6 of 18

TRANS
EN CRIBE 0.0 The quick brown ...

Next-Token
prediction
MLP
MLP
Self Attention
Cross Attention

Self Attention

Cross attention
Transformer
Encoder Blocks
MLP
MLP Transformer
Self Attention Decoder Blocks
Cross Attention

MLP Self Attention

Self Attention
MLP
Sinusoidal
Positional Cross Attention
Encoding
Self Attention

2x Conv 1D + GELU
Learned
Positional
Encoding

TRANS ...
SOT EN 0.0 the quick
CRIBE

º
Log-Mel Spectogram Tokens in Multitask Training Format

Figure 2. Whisper architecture representation. The log Mel-spectrograms are encoded by a Trans-
former network. Encoded representations are transformed into character outputs and no-speech
tokens via the Transformer decoder. Figure based on the one presented in [25].

The model was not fine-tuned in this study, thus the evaluation for all languages 194

was conducted in a zero-shot setting. The decoding was performed using a beam search 195

strategy with 5 beams, an array of temperature weights of [0.2, 0.4, 0.6, 0.8, 1], and a no 196

repeat n-gram size of 3 in order to take advantage of the language modeling head and to 197

avoid loops, in a similar way to [25]. 198

3. Materials 199

This section describes a set of open corpora used to benchmark the two considered 200

ASR systems (Section 3.1), followed by the performed process to derive a set of keywords to 201

be spotted by the considered systems (Section 3.2), and the description of a built in domain 202

dataset considered as well to test the considered models (Section 3.3). 203
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

7 of 18

3.1. Data 204

Table 1. List of public speech corpora considered to test the performance of ASR and KWS systems
based on Wav2Vec2.0 and Whisper.

Test Duration
Corpus name Description Languages
(h)
English 173
German 72
Read sentences collected French 38
Common Voice [38] and validated via Spanish 26
crowd-sourcing Italian 23
Portuguese 6
Polish 7
Spoken Wikipedia Volunteer readers of English 42
Corpus (SWC) [41] Wikipedia articles German 36
Speech segments from French 10
Media Speech [12]
YouTube videos Spanish 10
German 2
Audio recordings French 2
Multilingual TEDx [42] and transcripts Spanish 2
from TED talks Italian 2
Portuguese 2
Audio recordings
TEDLIUMv2 [43] English 3
from TED talks
German 14
French 10
Audio recordings Spanish 10
Multilingual librispeech (MLS) [37]
from audiobooks Italian 5
Polish 2
Portuguese 4
German 3
Crowdsourced French 4
Voxforge read Spanish 5
speech Italian 2
Portuguese 1
Audio recordings from
Debating technologies [44] English 1
transcribed public debates
Recordings from the
Polish Parliamentary corpus [45] Polish 1
Polish parliament
Combination of five
CORAA [20] Portuguese 13
corpora in Portuguese

Different public corpora were considered to train/test the ASR and KWS models. 205

Wav2vec2.0 models were fine-tuned using the Common Voice corpus [38] for each con- 206

sidered language. The amount of available labelled data highly varies depending on the 207

language, and include: 1600 hours for English, 777 hours for German, 623 for French, 324 208

for Spanish, 158 for Italian, 63 for Portuguese, and 43 for Polish. These data are freely 209

available via Huggingface4 . The training data for the Spanish model also included 57 hours 210

from the RTVE2018 dataset [46] from the Albayzin 2018 evaluation challenge. 211

The performance of both the fine-tuned Wav2vec2.0 and Whisper-based models was 212

evaluated in a cross-corpora fashion, considering a large set of databases from the literature 213

that are available in the different languages. The list of considered corpora is observed 214

4 [Link]
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

8 of 18

in Table 1. These corpora were selected in order to test the performance of the models in 215

several recording conditions, which can be closer to the realistic scenarios found by LEAs. 216

Notice that due to the sensitive nature of the target application, it is not possible to get 217

access to realistic operative data from LEAs. However, we created an in-domain synthetic 218

dataset using these open source corpora, which is described in Section 3.3. 219

3.2. Spotted keywords 220

In order to test the capabilities of the ASR models to spot specific keywords within 221

the child abuse domain, we defined a list of keywords to be spotted. The keyword list was 222

obtained from a set of open documents that include: (1) the "Best Practices on Victim support 223

for LEA first responders" deliverable from the GRACE project5 , (2) the 2021 "Barriers to 224

Compensation for Child Victims of Sexual Exploitation" report from ECPAT6 [47], (3) the 225

study from [48], (4) EUROPOL technical reports [49–51], (5) EUROPOL press-releases 226

from 2018 to 2022 using the keyword "child abuse"7 , (6) Wikipedia articles about "child 227

abuse" and "online child abuse", and (7) UNICEF press-releases about "child abuse"8 . 228

All documents were text crawled and pre-processed by performing lemmatisation, and 229

removing stop words, numbers, and date entities. After this process, we obtained a corpus 230

with 55,059 words, whose 6028 are unique. Figure 3 shows the most important keywords 231

found in the crawled corpus. 232

4
Corpus Presence (%)

0
ild

t
ab l
ex use

ontion
ma e
pa al
thr t

eo
b
oto

co net
pr on

r
mi e

rn ol
strphy
tra s
a
ua

gir
ren
ea

no

es
lin

um
we
i

o
iva
vid
ter

i
ch

sch
sex

erc
ph
er

ra
ita

int

og
plo

po

Words

Figure 3. Top 20 of the most important keywords related to child abuse, which were used to test the
capability of the ASR system to detect specific terminology within the domain.

Afterwards, we selected the 100 most repeated words from the corpus, which represent 233

the 33% of the information within the whole set of crawled documents. Finally, we excluded 234

12 terms because they were very broad concepts not related with child abuse, leading to a 235

final set of 88 keywords to be spotted. The obtained keyword list (in English) was translated 236

into the remaining six considered languages in order to have a common benchmark for all 237

languages. 238

3.3. GRACE dataset 239

We considered an additional corpus to test the implemented ASR systems by merging 240

and filtering the data described in Section 3.1. We selected audio samples from all datasets 241

5 [Link]
6 [Link]
7 [Link]
8 [Link]
5D=
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

9 of 18

that contain at least one of the 88 selected keywords. Table 2 shows the data distribution for 242

each language after selection. The table includes the datasets considered for each language 243

where the keywords are found, the number of utterances, and the total audio duration (in 244

hours). 245

Table 2. Data distribution for the GRACE dataset, which combines different corpora into a single one
within the child abuse domain.

Language Base corpora # utterances Duration (h)


English SWC, Debating technologies, TEDLIUMv2 2979 9.2
German Multilingual TEDx, SWC, Voxforge 1712 5.9
French Multilingual TEDx, MediaSpeech, Voxforge 1250 4.1
Spanish Multilingual TEDx, MediaSpeech, Voxforge 557 2.0
Italian Multilingual TEDx, Voxforge 354 1.0
Portuguese Multilingual TEDx, Voxforge, CORAA 1503 2.3

The selected audios were processed in order to have also more realistic acoustic 246

conditions than those expected in forensic applications within the considered domain. 247

The process includes: (1) adding background noise with signal to noise ratios (SNR) 248

between 5 and 30 dB (randomly), (2) adding reverberation using room impulses from the 249

VOiCES dataset [52], and (3) randomly applying the ogg-vorbis codec [53] due to it is 250

commonly found in audio material from online sources. The final ASR and KWS evaluation 251

is performed considering the two versions of the corpus: clean and noisy. This corpus is 252

available online9 to be used as a benchmark dataset for speech recognition in different 253

languages under non-controlled acoustic conditions. 254

4. Federated Learning 255

The considered FL pipeline is performed only with English data and includes five 256

nodes that are used for federated training, a dummy node considered to test the evolution 257

of the learning process, and the central server in charge of aggregating the weights received 258

from the five nodes. Figure 4 shows the implemented architecture. Three of the servers 259

were located at Vicomtech premises (Spain), one server was located at Greece, another one 260

in Portugal, and the remaining one in Cyprus. The aim of these connections is to create 261

a real environment for the pilot, in similar conditions to the expected when the model 262

is trained by different LEAs across Europe. In addition, secure communication between 263

clients and the server was established through a VPN connection to ensure that sensitive 264

data (parameters) are safely transmitted and to prevent unauthorised access. Each node 265

contains data from a different dataset: TEDLIUMv2, debating technologies, Librispeech- 266

other, Librispeech-clean, and SWC. This data configuration aims to evaluate the impact of 267

non-IID data distribution, which is more realistic for the addressed forensic application. 268

9 [Link]
GRACE_ASR.zip
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

10 of 18

site-1
TEDLIUM v2

site-2
site-5 Debating
SWC technologies

local m odel aggr egated


update m odel update

aggr egated local m odel aggr egated


local m odel
m odel update m odel
update
update update

local m odel local m odel


site-4 update update site-3
Libr ispeech clean Libr ispeech other

aggr egated aggr egated


m odel update m odel update

dum my node
Libr ispeech test clean

Figure 4. Configuration of the FL architecture. Central server with five client nodes (site-{1, 2, · · · , 5})
and a dummy node only used to test the performance of the aggregated model

The FL pilot was performed only with the Wav2Vec2.0 model, and using also the 269

pre-trained Wav2Vec2-XLS-R-300M model. The training hyperparameters were the same 270

for the five clients, and include a batch size of 2, a learning rate of 5 × 10−5 warmed up 271

in the first 10% of the training time, and a gradient accumulation of 16 steps. The local 272

training is performed for 5 epochs. The central server is configured to run for 10 rounds of 273

federated training, and using the Federated averaging (FedAvg) aggregation mechanism 274

to update the central model. The architecture configuration and the training process is 275

implemented using Nvidia Flare10 . 276

5. Results 277

5.1. Speech Recognition 278

Wav2Vec2.0 and Whisper models were evaluated under the described corpora in 279

Section 3.1. The results of the ASR systems in terms of WER are shown in Table 3. The 280

results included those obtained in the evaluation of the seven languages, and using both the 281

open benchmark corpora and the two versions (clean and noisy) of the synthetic GRACE 282

corpus. 283

10 [Link]
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

11 of 18

Table 3. Results of the ASR models in different languages considering all benchmark datasets. Results
in terms of WER.

Model Common MLS TED- MTEDx SWC Media Voxforge Debates Polish CORAA GRACE GRACE AVG.
Voice LIUMv2 Speech Parl clean noisy
English
Wav2Vec 2.0 16.1 - 17.2 - 20.6 - - 11.7 - - 18.9 32.6 19.5
Whisper 10.0 - 5.4 - 20.6 - - 7.0 - - 24.5 19.8 14.6
German
Wav2Vec 2.0 11.9 12.9 - 36.7 34.5 - 7.5 - - - 20.0 33.8 22.5
Whisper 7.1 6.7 - 21.7 18.3 4.2 - - - 15.5 22.5 13.7
French
Wav2Vec 2.0 16.7 17.0 - 25.3 - 29.1 16.7 - - - 26.5 56.3 26.8
Whisper 21.7 8.0 - 23.3 - 35.8 14.6 - - - 36.8 34.1 24.9
Spanish
Wav2Vec 2.0 4.7 7.2 - 12.9 - 14.5 6.3 - - - 12.6 33.3 13.1
Whisper 6.2 5.3 - 9.4 - 15.8 4.2 - - - 19.6 18.8 11.3
Italian
Wav2Vec 2.0 12.8 21.1 - 22.2 - - 14.3 - - - 18.0 46.3 22.5
Whisper 7.9 13.6 - 11.6 - - 10.5 - - - 14.1 20.2 13.0
Portuguese
Wav2Vec 2.0 12.9 20.1 - 33.8 - - 17.8 - - 48.5 42.7 68.1 34.8
Whisper 5.4 8.8 - 13.1 - - 11.2 - - 21.7 22.1 42.3 17.8
Polish
Wav2Vec 2.0 11.5 12.7 - - - - - - 32.1 - - - 18.8
Whisper 8.9 6.0 - - - - - - 32.5 - - - 15.8

On average, the WER for each language using Whisper ranges from 11.3% (in Spanish) 284

to 24.9% (in French). The results using Wav2Vec2.0 range from 13.1% (in Spanish) to 34.8% 285

(in Portuguese). In general, Whisper produces less errors than Wav2Vec2.0 (see Figure 5 286

left). The difference between both models is statistically significant according to a Mann 287

Whitney test (U=1203.5, p-value=0.016). Whisper outperformed Wav2Vec2.0 especially 288

under the most affected acoustic conditions, such as in the GRACE noisy, TEDLIUMv2, 289

Debates, and CORAA corpora. However, there are some scenarios where Wav2Vec2.0 290

outperformed Whisper and which should be considered with special attention, such as the 291

results for Spanish Common Voice. 292

The results obtained were compared to those found in the literature for the multilingual 293

corpora: Common Voice, MLS, MTEDx, and MediaSpeech. The comparison is shown in 294

Table 4. The Wav2Vec2.0-based model outperformed results in the Spanish versions of 295

Common Voice and MediaSpeech corpora, with WERs of 4.3% and 14.5%, respectively 296

with respect to to the results reported in [18] for Common Voice (WER=6.2%) and in [12] 297

for MediaSpeech (WER=18.3%). We also reported state-of-the-art results for the Spanish, 298

Portuguese, Italian, and German versions of the MTEDx corpus (WERs of 9.4%, 12%, 299

11.6%, and 21.7%, respectively) with respect to the WERs of 16.2%, 20.2%, 16.4%, and 300

42.3% reported in [42]. Whisper model also achieved state-of-the-art results in the CORAA 301

corpus (WER=21.7%) with respect to the results reported in [54] (WER=21.9%), and in 302

the TEDLIUMv2 corpus (WER=5.4%) compared to [13] (WER=5.6%). Regarding MLS, the 303

state-of-the-art results are still from [55]. However, notice that the results reported here 304

correspond to cross-corpus tests, while the experiments performed in [55] correspond to 305

Wav2Vec2.0 models trained and tested using MLS, thus making the models adapted just 306

for such a corpus. 307


Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

12 of 18

Table 4. WER comparison between the results reported and those coming from the state-of-the-art for
Common Voice, MLS, and MTEDx corpora. Best results for each corpus and language are highlighted
in bold

Corpus Reference Language


English German French Spanish Italian Portuguese Polish
[11] - 7.7 12.5 10.9 - - -
[18] - 7.2 11.2 6.2 6.5 6.1 7.6
[36] - 7.8 9.6 10.0 - - -
[25] 10.1 7.7 14.7 6.4 8.1 7.1 9.0
Common [19] - 3.6 - - - - -
Voice [56] - 9.8 - - - - -
[57] - - - - - 9.2 -
Wav2vec2.0 16.1 11.9 16.7 4.7 12.8 12.9 11.5
Whisper-large 10.0 7.1 21.7 6.2 7.9 5.4 8.9
[37] - 6.5 5.6 6.1 10.5 19.5 20.4
[40] - 7.4 10.0 6.9 12.0 15.6 9.8
[55] - 4.1 5.0 3.7 8.2 8.0 6.6
MLS [25] - 6.6 8.9 5.4 14.3 9.2 6.6
[57] - - - - - 12.3 -
Wav2vec2.0 - 12.9 17.0 7.2 21.1 20.1 12.7
Whisper-large - 6.7 8.0 5.3 15.8 8.8 9.9
[42] - 42.3 19.4 16.2 16.4 20.2 -
[57] - - - - - 21.0 -
MTEDx
Wav2vec2.0 - 36.7 25.3 12.9 22.2 33.8 -
Whisper-large - 21.7 23.3 9.4 11.6 13.1 -
[12] - 19.2 18.3 - - - -
MediaSpeech Wav2vec2.0 - 29.1 14.5 - - - -
Whisper-large - 35.8 15.8 - - - -

5.2. Keyword Spotting 308

The text transcriptions from Wav2Vec2.0 and Whisper were post-processed in order to 309

find the presence of the defined keywords to be spotted. The process involved transforming 310

the transcription to lowercase and lemmatization. Lemmatization is performed to reduce 311

the inflectional form of each word in order to detect all possible variations of the word 312

within the transcription. The lemmatization process is performed using the set of large 313

open dictionaries available in Spacy11 . The results obtained for KWS in each corpus are 314

shown in Table 5. The results are presented in terms of the true positive rate (TPR). This is 315

a common metric used in this type of applications where it is more important to avoid false 316

positive than false negative errors [58,59]. 317

On average, the TPRs are higher using Whisper, and the results per language using 318

Whisper range from 81.5% (for Polish) to 98.4% (for Italian). Results using Wav2Vec2.0 319

range from 82.9% (for Portuguese) to 94.9% (for Spanish). Similar to the ASR results, the 320

difference between Whisper and Wav2Vec2.0 is larger when considering speech signals 321

in non-controlled acoustic conditions, like the ones from the GRACE noisy corpus, where 322

we particularly guarantee the presence of the spotted keywords in every utterance. High 323

differences were also observed int the CORAA corpus, in Common Voice, and in the 324

German SWC. The difference between the results obtained using Wav2Vec2.0 and Whisper 325

is also statistically significant (see Figure 5 right) according to a Mann-Whitney test with 326

U=589.0 and a p-value=0.003. 327

11 [Link]
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

13 of 18

Table 5. Results of KWS in different languages considering all benchmark datasets. Results in terms
of TPR (%).

Model Common MLS TED- MTEDx SWC Media Voxforge Debates Polish CORAA GRACE GRACE AVG.
Voice LIUMv2 Speech Parl clean noisy
English
Wav2Vec 2.0 93.3 - 95.4 - 92.5 - - 96.6 - - 94.6 79.5 92.0
Whisper 96.8 - 97.4 - 94.1 - - 97.7 - - 91.8 93.6 95.2
German
Wav2Vec 2.0 91.3 96.9 - 93.5 80.4 - 99.8 - - - 94.3 79.7 90.8
Whisper 97.8 98.8 - 97.7 97.6 99.8 - - - 97.2 90.6 97.1
French
Wav2Vec 2.0 90.6 90.1 - 94.5 - 82.8 90.7 - - - 90.5 60.1 85.6
Whisper 94.6 98.0 - 93.9 - 84.9 94.2 - - - 88.0 81.3 90.7
Spanish
Wav2Vec 2.0 96.7 98.1 - 98.1 - 96.3 100.0 - - - 97.0 78.1 94.9
Whisper 98.2 99.8 - 98.6 - 94.4 99.8 - - - 92.0 94.2 96.7
Italian
Wav2Vec 2.0 90.7 97.2 - 95.5 - - 98.9 - - - 96.1 80.9 93.2
Whisper 97.8 99.9 - 97.3 - - 99.8 - - - 98.7 96.9 98.4
Portuguese
Wav2Vec 2.0 93.1 93.4 - 94.1 - - 99.1 - - 74.3 76.8 49.5 82.9
Whisper 96.6 97.5 - 99.4 - - 100.0 - - 88.1 88.3 81.3 93.0
Polish
Wav2Vec 2.0 93.9 96.9 - - - - - - 83.3 - - - 91.4
Whisper 95.4 98.7 - - - - - - 50.3 - - - 81.5

0.06 Mann-Whitney U: 1203.5, p-value: 0.016


Wav2Vec 2.0 0.175 Mann-Whitney U: 589.0, p-value: 0.003
Whisper
0.05 0.150

0.04 0.125
Normalized count
Normalized count

0.100
0.03
0.075
0.02
0.050
0.01 0.025
0.00 0.000
0 10 20 30 40 50 60 70 80 30 40 50 60 70 80 90 100
WER (%) TPR (%)

Figure 5. Comparison between the results obtained using Wav2Vec2.0 and Whisper for ASR (left)
and KWS (right).

5.3. Federated Learning 328

The FL experiment involved training the Wav2Vec2.0 system using 5 separate real 329

servers for training, and one additional node (dummy) used only to test the final model. 330

Each node contained data from a different dataset (only in English) in order to evaluate 331

the contribution from each corpus into the global aggregated model. The aim was also to 332

cover non-IID conditions, which have shown to be one of the most important drawbacks 333

when training models in an FL approach. The results are shown in Table 6. The results 334

using the FL training are compared to those obtained training the system in a complete 335

centralized manner. Similar WERs were obtained in each node comparing the federated 336

vs. centralized training. The main difference is that when considering FL models there is 337

only one aggregated model which covers the results of the 5 nodes, instead of having 5 338

different models for the case of the centralized approach. This fact highly reduces the time 339

considered to train the system, and most important, it is possible to take advantage of data 340
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

14 of 18

from different data centers to train a more robust and general model without the need of 341

sharing data among clients. 342

Table 6. Results of the FL pilot comparing WERs from Wav2Vec2.0 models trained in a federated or
centralized way.

Node Data WER Federated WER centralized


node-1 TED-LIUMv2 13.5 13.3
node-2 Debates 12.4 12.3
node-3 Librispeech-other 7.9 7.8
node-4 Librispeech-clean 2.8 3.2
node-5 SWC 25.7 24.3
dummy Librispeech-clean 3.6 3.8

6. Discussion 343

The evaluation of Wav2Vec2.0 and Whisper-based ASR systems is performed under a 344

large set of different scenarios, including one specifically designed for forensic applications 345

within child domain abuse. On average, Whisper is more accurate than the Wav2Vec2.0- 346

based system. Whisper achieved WERs ranging from 11.5% to 24.9%, depending on the 347

language, compared with Wav2Vec2.0 WERs that range between 13.3% to 34.8%. The 348

difference between the two models is even larger when considering languages trained with 349

lower resources, such as Portuguese or Italian. Although these differences, Wav2vec2.0 350

is competitive with Whisper when the number of hours for fine-tuning is large, e.g, in 351

English, Spanish, or French. 352

Results using the GRACE dataset show relatively similar WERs between Wav2Vec2.0 353

and Whisper when considering the clean version of the corpus, with an average WER of 354

22.1% for Whisper and of 23.2% for Wav2Vec2.0. However, the difference between the 355

two models greatly increases when considering the noisy version of the corpus, with an 356

average WER of 26.3% for Whisper and of 45.1% for Wav2Vec2.0. This is a great indicator 357

about the capability of Whisper to perform accurate transcriptions under non-controlled 358

and noisy acoustic conditions, by keeping similar WERs in the two versions of the GRACE 359

corpus. Despite the differences between the two types of models, there are some surprising 360

results where Wav2Vec2.0 outperforms Whisper, and which should be considered with 361

special attention. For instance, when evaluating the GRACE clean corpus in languages such 362

as English, French, and Spanish. The models for these three languages were fine-tuned 363

with more data, which likely explains the WER reduction in Wav2Vec2.0 with respect to 364

Whisper. 365

The performed evaluations of our systems achieved state-of-the art results in several 366

of the considered benchmark corpora. We reported state-of-the-art results for some of the 367

languages in the Common Voice corpus. State-of-the-art results were also achieved for 368

almost all languages in the MTEDx and MediaSpeech corpora. These results are good 369

indicators about the capabilities of the considered systems to accurately recognize speech 370

under more natural and spontaneous scenarios, closer to the expected in forensic domains. 371

The KWS evaluation indicated that both Wav2Vec2.0 and Whisper were accurate 372

enough to recognize the considered child abuse-related keywords in the seven languages. 373

TPRs obtained for Wav2Vec2.0 range from 82.9% to 94.9%, depending on the language. 374

Results using Whisper range from 80.3% to 98.2%. The particular evaluation of KWS in the 375

GRACE dataset also shows that both models are equally accurate to recognize the selected 376

keywords under controlled acoustic conditions. On the contrary, when considering the 377

noisy version of the corpus, the results for Wav2Vec2.0 are reduced by 20% while the results 378

for Whisper are only reduced by 3%. This fact again indicates the capability of Whisper to 379

accurately process speech recordings in non-controlled acoustic conditions. 380

The last covered experiment involved a pilot study on the use of FL to train ASR 381

systems. The results indicated that an ASR trained in a federated way maintains and 382

in some cases outperforms the performance of individual ASRs trained in a centralized 383
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

15 of 18

manner by each LEA. In addition to the performance, the most important aspect of FL is 384

that the ASR training does not involve any data sharing among LEAs, since only updates 385

of the network parameters are transferred to a central server in charge of aggregating the 386

model. These results are indicators about the potential use of FL to obtain a joint (and 387

potentially richer) model combining sources of data that could not be otherwise combined. 388

Although the benefits of using FL, it is important to consider external factors that may 389

degrade the performance and reliability of the system. For instance, there is evidence about 390

FL attacks that are able to retrieve speaker information from the transferred weights [60] 391

or data poisoning attacks inside LEAs server. Different strategies can be considered to 392

mitigate this this type of attacks such as the use of differential privacy algorithms [61] or 393

the use of trusted execution environments. 394

7. Conclusions 395

This paper proposed the use of speech recognition and keyword spotting technologies 396

to be applied in forensic scenarios, particularly in child exploitation domains. The aim is 397

to provide LEAs with technology to detect the presence of offensive online audiovisual 398

material related to child abuse. State-of-the art ASR systems based on Wav2Vec2.0 and 399

Whisper were considered for the addressed application. The performance of both models 400

was tested on a large set of open benchmark corpora from the literature. Therefore, the 401

results obtained can be extended to other ASR domains. We additionally created an in- 402

domain corpus using different open source datasets from the research community. The aim 403

was to test the models in more realistic and operative conditions. 404

The ASR and KWS models were evaluated in corpora from seven Indo-European 405

languages, including English, German, French, Spanish, Italian, Portuguese, and Polish. 406

We obtained overall WERs ranging from 11.3% to 24.9%, depending on the language. The 407

performance of the KWS model for the different languages ranged from 81.5% to 98.4%. 408

The most accurate results were obtained from models trained with more data, such as 409

English or German. The comparison between Wav2Vec2.0 and Whisper models indicated 410

that the second one was the most accurate system in the majority of cases, especially when 411

considering utterances in non-controlled acoustic conditions. 412

We also proposed a strategy for using FL to train robust ASR systems in the context 413

of the addressed application. This is a suitable approach considering that collecting op- 414

erational data from LEAs is not possible. FL approaches allow LEAs to build a common 415

technological platform without the need to share their operational data. The results of the 416

FL pilot indicated that similar WERs were achieved when comparing the model trained in 417

a federated way to individual models trained in a centralized manner, even considering 418

non-IID conditions, which has been shown to be one of the main drawbacks in FL. 419

For future work, the considered approaches can be extended to other forensic applica- 420

tions where there is a need to monitor audiovisual material from online sources. In addition, 421

the considered technology can be combined with other speech processing methods, such as 422

speaker and language identification, age and gender recognition, and speaker diarization. 423

The ultimate goal is to provide LEAs with accurate tools to monitor audio from online 424

sources, allowing them to respond in a practical and timely manner. 425

Author Contributions: "Conceptualization, J.C.V., A.A., and .; methodology, J.C.V. and A.A; software, 426

J.C.V.; validation, J.C.V.; formal analysis, J.C.V. and A.A; investigation, J.C.V, A.A., and .; resources, .; 427

data curation, J.C.V.; writing—original draft preparation, J.C.V.; writing—review and editing, J.C.V., 428

A.A., and .; visualization, J.C.V. All authors have read and agreed to the published version of the 429

manuscript.” 430

Funding: This project has received funding from the European Union’s Horizon 2020 research and 431

innovation programme under project GRACE, grant agreement No 883341 432

Institutional Review Board Statement: The study was conducted in accordance with the Declaration 433

of Helsinki, and approved by the Institutional Review Board of the GRACE consortium (protocol 434

code XXX and date of approval). 435


Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

16 of 18

Data Availability Statement: All data considered in this study come from open repositories under 436

Creative common licenses. 437

Conflicts of Interest: "The authors declare no conflict of interest.” 438

Abbreviations 439

ASR Automatic Speech Recognition


CTC Connectionist Temporal Classification
FL Federated Learning
KWS Keyword Spotting
440
LEA Law Enforcement Agency
MLS Multilingual Librispeech
SWC Spoken Wikipedia Corpus
TPR True Positive Rate

References 441

1. Negrão, M.; Domingues, P. SpeechToText: An open-source software for automatic detection and transcription of voice recordings 442

in digital forensics. Forensic Science International: Digital Investigation 2021, 38, 301223. 443

2. Alghowinem, S. A safer youtube kids: An extra layer of content filtering using automated multimodal analysis. In Proceedings 444

of the Proceedings of SAI Intelligent Systems Conference. Springer, 2018, pp. 294–308. 445

3. Mariconti, E.; Suarez-Tangil, G.; Blackburn, J.; De Cristofaro, E.; Kourtellis, N.; Leontiadis, I.; Serrano, J.L.; Stringhini, G. " You 446

Know What to Do" Proactive Detection of YouTube Videos Targeted by Coordinated Hate Attacks. Proceedings of the ACM on 447

Human-Computer Interaction 2019, 3, 1–21. 448

4. Amodei, D.; Ananthanarayanan, S.; Anubhai, R.; Bai, J.; Battenberg, E.; Case, C.; Casper, J.; Catanzaro, B.; Cheng, Q.; Chen, G.; 449

et al. Deep speech 2: End-to-end speech recognition in english and mandarin. In Proceedings of the International conference on 450

machine learning. PMLR, 2016, pp. 173–182. 451

5. Graves, A.; Jaitly, N. Towards end-to-end speech recognition with recurrent neural networks. In Proceedings of the International 452

conference on machine learning. PMLR, 2014, pp. 1764–1772. 453

6. Chan, W.; Jaitly, N.; Le, Q.; Vinyals, O. Listen, attend and spell: A neural network for large vocabulary conversational speech 454

recognition. In Proceedings of the ICASSP. IEEE, 2016, pp. 4960–4964. 455

7. Chorowski, J.K.; Bahdanau, D.; Serdyuk, D.; Cho, K.; Bengio, Y. Attention-based models for speech recognition. Advances in 456

neural information processing systems 2015, 28. 457

8. Lu, L.; Zhang, X.; Renais, S. On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech 458

recognition. In Proceedings of the ICASSP. IEEE, 2016, pp. 5060–5064. 459

9. Yao, Z.; Wu, D.; Wang, X.; Zhang, B.; Yu, F.; Yang, C.; Peng, Z.; Chen, X.; Xie, L.; Lei, X. Wenet: Production oriented streaming and 460

non-streaming end-to-end speech recognition toolkit. arXiv preprint arXiv:2102.01547 2021. 461

10. Kriman, S.; Beliaev, S.; Ginsburg, B.; Huang, J.; Kuchaiev, O.; Lavrukhin, V.; Leary, R.; Li, J.; Zhang, Y. Quartznet: Deep automatic 462

speech recognition with 1d time-channel separable convolutions. In Proceedings of the ICASSP. IEEE, 2020, pp. 6124–6128. 463

11. Bermuth, D.; Poeppel, A.; Reif, W. Scribosermo: Fast Speech-to-Text models for German and other Languages. arXiv preprint 464

arXiv:2110.07982 2021. 465

12. Kolobov, R.; Okhapkina, O.; Omelchishina, O.; Platunov, A.; Bedyakin, R.; Moshkin, V.; Menshikov, D.; Mikhaylovskiy, N. 466

Mediaspeech: Multilanguage asr benchmark and dataset. arXiv preprint arXiv:2103.16193 2021. 467

13. Majumdar, S.; Balam, J.; Hrinchuk, O.; Lavrukhin, V.; Noroozi, V.; Ginsburg, B. Citrinet: Closing the gap between non- 468

autoregressive and autoregressive end-to-end models for automatic speech recognition. arXiv preprint arXiv:2104.01721 2021. 469

14. Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the Proceedings of the IEEE conference on computer 470

vision and pattern recognition, 2018, pp. 7132–7141. 471

15. Graves, A. Sequence transduction with recurrent neural networks. arXiv preprint arXiv:1211.3711 2012. 472

16. Zhou, W.; Zheng, Z.; Schlüter, R.; Ney, H. On language model integration for rnn transducer based speech recognition. In 473

Proceedings of the ICASSP. IEEE, 2022, pp. 8407–8411. 474

17. Baevski, A.; et al. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information 475

Processing Sysrecognitiontems 2020, 33, 12449–12460. 476

18. Pham, N.Q.; Waibel, A.; Niehues, J. Adaptive multilingual speech recognition with pretrained models. In Proceedings of the 477

INTERSPEECH, 2022, pp. 3879–3883. [Link] 478

19. Krabbenhöft, H.N.; Barth, E. TEVR: Improving Speech Recognition by Token Entropy Variance Reduction. arXiv preprint 479

arXiv:2206.12693 2022. 480

20. Junior, A.C.; Casanova, E.; Soares, A.; de Oliveira, F.S.; Oliveira, L.; Junior, R.C.F.; da Silva, D.P.P.; Fayet, F.G.; Carlotto, B.B.; Gris, 481

L.R.S.; et al. CORAA: a large corpus of spontaneous and prepared speech manually validated for speech recognition in Brazilian 482

Portuguese. arXiv preprint arXiv:2110.15731 2021. 483


Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

17 of 18

21. Hsu, W.N.; Tsai, Y.H.H.; Bolte, B.; Salakhutdinov, R.; Mohamed, A. HuBERT: How much can a bad teacher benefit ASR 484

pre-training? In Proceedings of the ICASSP. IEEE, 2021, pp. 6533–6537. 485

22. Hsu, W.N.; Bolte, B.; Tsai, Y.H.H.; Lakhotia, K.; Salakhutdinov, R.; Mohamed, A. Hubert: Self-supervised speech representation 486

learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing 2021, 29, 3451– 487

3460. 488

23. Gulati, A.; Qin, J.; Chiu, C.C.; Parmar, N.; Zhang, Y.; Yu, J.; Han, W.; Wang, S.; Zhang, Z.; Wu, Y.; et al. Conformer: Convolution- 489

augmented Transformer for Speech Recognition. In Proceedings of the Proc. Interspeech 2020, 2020, pp. 5036–5040. https: 490

//[Link]/10.21437/Interspeech.2020-3015. 491

24. Guo, P.; Boyer, F.; Chang, X.; Hayashi, T.; Higuchi, Y.; Inaguma, H.; Kamo, N.; Li, C.; Garcia-Romero, D.; Shi, J.; et al. Recent 492

developments on espnet toolkit boosted by conformer. In Proceedings of the ICASSP. IEEE, 2021, pp. 5874–5878. 493

25. Radford, A.; Kim, J.W.; Xu, T.; Brockman, G.; McLeavey, C.; Sutskever, I. Robust speech recognition via large-scale weak 494

supervision. Technical report, OpenAI, 2022. 495

26. Voigt, P.; Von dem Bussche, A. The eu general data protection regulation (gdpr). A Practical Guide, 1st Ed., Cham: Springer 496

International Publishing 2017, 10, 10–5555. 497

27. Konečnỳ, J.; McMahan, H.B.; Yu, F.X.; Richtárik, P.; Suresh, A.T.; Bacon, D. Federated learning: Strategies for improving 498

communication efficiency. arXiv preprint arXiv:1610.05492 2016. 499

28. Yang, Q.; Liu, Y.; Cheng, Y.; Kang, Y.; Chen, T.; Yu, H. Federated learning. Synthesis Lectures on Artificial Intelligence and Machine 500

Learning 2019, 13, 1–207. 501

29. Li, L.; Fan, Y.; Tse, M.; Lin, K.Y. A review of applications in federated learning. Computers & Industrial Engineering 2020, 502

149, 106854. 503

30. Li, T.; Sahu, A.K.; Talwalkar, A.; Smith, V. Federated learning: Challenges, methods, and future directions. IEEE Signal Processing 504

Magazine 2020, 37, 50–60. 505

31. Dimitriadis, D.; Kumatani, K.; Gmyr, R.; Gaur, Y.; Eskimez, S.E. A Federated Approach in Training Acoustic Models. In 506

Proceedings of the Interspeech, 2020, pp. 981–985. 507

32. Cui, X.; Lu, S.; Kingsbury, B. Federated acoustic modeling for automatic speech recognition. In Proceedings of the ICASSP. IEEE, 508

2021, pp. 6748–6752. 509

33. Guliani, D.; Beaufays, F.; Motta, G. Training speech recognition models with federated learning: A quality/cost framework. In 510

Proceedings of the ICASSP. IEEE, 2021, pp. 3080–3084. 511

34. Hard, A.; Partridge, K.; Nguyen, C.; Subrahmanya, N.; Shah, A.; Zhu, P.; Moreno, I.L.; Mathews, R. Training Keyword 512

Spotting Models on Non-IID Data with Federated Learning. In Proceedings of the INTERSPEECH, 2020, pp. 4343–4347. 513

[Link] 514

35. Conneau, A.; Baevski, A.; Collobert, R.; Mohamed, A.; Auli, M. Unsupervised cross-lingual representation learning for speech 515

recognition. arXiv preprint arXiv:2006.13979 2020. 516

36. Wang, C.; Riviere, M.; Lee, A.; Wu, A.; Talnikar, C.; Haziza, D.; Williamson, M.; Pino, J.; Dupoux, E. VoxPopuli: A Large-Scale 517

Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation. In Proceedings of 518

the Annual Meeting of the Association for Computational Linguistics and International Joint Conference on Natural Language 519

Processing (Volume 1: Long Papers), 2021, pp. 993–1003. 520

37. Pratap, V.; Xu, Q.; Sriram, A.; Synnaeve, G.; Collobert, R. MLS: A Large-Scale Multilingual Dataset for Speech Research. In 521

Proceedings of the INTERSPEECH, 2020. 522

38. Ardila, R.; Branson, M.; Davis, K.; Henretty, M.; Kohler, M.; Meyer, J.; Morais, R.; Saunders, L.; Tyers, F.M.; Weber, G. Common 523

Voice: A Massively-Multilingual Speech Corpus. In Proceedings of the Proceedings of the 12th Conference on Language 524

Resources and Evaluation (LREC 2020), 2020, pp. 4211–4215. 525

39. Valk, J.; Alumäe, T. VoxLingua107: a dataset for spoken language recognition. In Proceedings of the IEEE Spoken Language 526

Technology Workshop (SLT). IEEE, 2021, pp. 652–658. 527

40. Babu, A.; Wang, C.; Tjandra, A.; Lakhotia, K.; Xu, Q.; Goyal, N.; Singh, K.; von Platen, P.; Saraf, Y.; Pino, J.; et al. XLS-R: Self- 528

supervised Cross-lingual Speech Representation Learning at Scale. In Proceedings of the INTERSPEECH, 2022, pp. 2278–2282. 529

[Link] 530

41. Baumann, T.; Köhn, A.; Hennig, F. The Spoken Wikipedia Corpus collection: Harvesting, alignment and an application to 531

hyperlistening. Language Resources and Evaluation 2019, 53, 303–329. 532

42. Salesky, E.; Wiesner, M.; Bremerman, J.; Cattoni, R.; Negri, M.; Turchi, M.; Oard, D.W.; Post, M. The Multilingual TEDx Corpus 533

for Speech Recognition and Translation. In Proceedings of the INTERSPEECH, 2021, pp. 3655–3659. [Link] 534

/Interspeech.2021-11. 535

43. Rousseau, A.; Deléglise, P.; Esteve, Y.; et al. Enhancing the TED-LIUM corpus with selected data for language modeling and more 536

TED talks. In Proceedings of the LREC, 2014, pp. 3935–3939. 537

44. Mirkin, S.; Jacovi, M.; Lavee, T.; Kuo, H.K.; Thomas, S.; Sager, L.; Kotlerman, L.; Venezian, E.; Slonim, N. A Recorded Debating 538

Dataset. In Proceedings of the LREC, 2017, pp. 250–254. 539

45. Ogrodniczuk, M. Polish parliamentary corpus. In Proceedings of the LREC, 2018, pp. 15–19. 540

46. Lleida, E.; Ortega, A.; Miguel, A.; Bazán-Gil, V.; Pérez, C.; Gómez, M.; De Prada, A. Albayzin 2018 evaluation: the iberspeech-rtve 541

challenge on speech technologies for spanish broadcast media. Applied sciences 2019, 9, 5412. 542
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1

18 of 18

47. ECPAT. Barriers to Compensation for Child Victims of Sexual Exploitation A discussion paper based on a comparative legal 543

study of selected countries. ECPAT Internaltional 2021. 544

48. Richards, K. Misperceptions about child sex offenders. Trends and issues in crime and criminal justice 2011, pp. 1–8. 545

49. EUROPOL. Online sexual coercion and extortion as a form of crime affecting children. European Union Agency for Law Enforcement 546

Cooperation 2017. 547

50. EUROPOL. Internet Organised Crime Threat Assessment. European Union Agency for Law Enforcement Cooperation 2019. 548

51. EUROPOL. Exploting Isolation: Offenders and victims of online child sexual abuse during the COVID-19 pandemic. European 549

Union Agency for Law Enforcement Cooperation 2020. 550

52. Richey, C.; Barrios, M.A.; Armstrong, Z.; Bartels, C.; Franco, H.; Graciarena, M.; Lawson, A.; Nandwana, M.K.; Stauffer, A.; van 551

Hout, J.; et al. Voices Obscured in Complex Environmental Settings (VOiCES) Corpus. In Proceedings of the INTERSPEECH, 552

2018, pp. 1566–1570. [Link] 553

53. Moffitt, J. Ogg Vorbis—open, free audio—set your media free. Linux journal 2001, 2001, 9–es. 554

54. Marcacini, R.M.; Candido Junior, A.; Casanova, E. Overview of the Automatic Speech Recognition for Spontaneous and Prepared 555

Speech & Speech Emotion Recognition in Portuguese (SE&R) Shared-tasks at PROPOR 2022. In Proceedings of the PROPOR, 556

2022. 557

55. Bai, J.; Li, B.; Zhang, Y.; Bapna, A.; Siddhartha, N.; Sim, K.C.; Sainath, T.N. Joint unsupervised and supervised training for 558

multilingual asr. In Proceedings of the ICASSP. IEEE, 2022, pp. 6402–6406. 559

56. Zheng, H.; Peng, W.; Ou, Z.; Zhang, J. Advancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces and 560

Conformers. arXiv preprint arXiv:2107.03007 2021. 561

57. Stefanel Gris, L.R.; Casanova, E.; Oliveira, F.S.d.; Silva Soares, A.d.; Candido Junior, A. Brazilian Portuguese Speech Recognition 562

Using Wav2vec 2.0. In Proceedings of the International Conference on Computational Processing of the Portuguese Language. 563

Springer, 2022, pp. 333–343. 564

58. Keshet, J.; Grangier, D.; Bengio, S. Discriminative keyword spotting. Speech Communication 2009, 51, 317–329. 565

59. Lengerich, C.; Hannun, A. An end-to-end architecture for keyword spotting and voice activity detection. arXiv preprint 566

arXiv:1611.09405 2016. 567

60. Tomashenko, N.; Mdhaffar, S.; Tommasi, M.; Estève, Y.; Bonastre, J.F. Privacy attacks for automatic speech recognition acoustic 568

models in a federated learning framework. In Proceedings of the ICASSP. IEEE, 2022, pp. 6972–6976. 569

61. Geyer, R.C.; Klein, T.; Nabi, M. Differentially private federated learning: A client level perspective. arXiv preprint arXiv:1712.07557 570

2017. 571

Common questions

Powered by AI

Federated Learning has the potential to significantly impact the collaborative development of ASR technology among law enforcement agencies by enabling shared model improvements without compromising data security. It facilitates the combination of models trained on diverse datasets, leading to richer, more robust systems. This collaboration can yield models that are better suited to various dialects and conditions met by different agencies while maintaining strict data privacy, thus enhancing ASR's applicability in widespread forensic operations .

Both Whisper and Wav2Vec2.0 are effective in recognizing child abuse-related keywords across different languages, achieving significant True Positive Rates (TPRs). Wav2Vec2.0's TPRs range from 82.9% to 94.9%, while Whisper's range is 80.3% to 98.2%. These results underscore Whisper's slightly superior performance in handling various languages, likely due to its large training dataset and robust handling of noise .

Federated Learning (FL) addresses privacy concerns by decentralizing the training process, allowing models to be trained on local devices without the need to share sensitive data. Instead of sending raw data, only updates of the network parameters are transferred to a central server, which aggregates these updates to improve a global model. This method keeps the data on-premise, ensuring that sensitive information relating to law enforcement activities is not transmitted elsewhere, thereby maintaining privacy and data security .

Federated Learning (FL) enhances the robustness of ASR models by allowing them to be trained on diverse datasets from multiple sources without centralizing the actual data. This approach helps overcome data heterogeneity issues, as insights and patterns from various local datasets are aggregated into a central model, providing a more generalized model without compromising data privacy. FL has been shown to maintain, and in some cases exceed, the performance of models trained centrally .

Whisper is considered more suitable for forensic applications primarily due to its superior performance under noisy and non-controlled environments, which are common in forensic settings. The model achieves lower WERs in these challenging acoustic environments, maintaining a consistent level of accuracy where Wav2Vec2.0 performance falls considerably .

The Whisper model demonstrates a significant advantage in noisy acoustic conditions compared to Wav2Vec2.0. When tested on the GRACE dataset, the average WER for Whisper in noisy conditions was 26.3% compared to 45.1% for Wav2Vec2.0. This showcases Whisper's strong capability to maintain accuracy in non-controlled environments, reducing the impact of noise substantially less than Wav2Vec2.0 does .

Data heterogeneity in federated training poses a challenge as it can lead to inefficiencies and degraded model performance if the data distribution varies significantly across clients. To address this, researchers have implemented client adaptive federated training strategies and random client data sampling methods. These approaches aim to balance the quality-cost trade-off and compensate for non-IID data, thereby achieving comparable WERs to fully centralized systems .

When designing ASR solutions for forensic scenarios, the main considerations include selecting neural architectures that can handle various acoustic environments and ensuring data privacy and protection. These considerations lead to the preference for models like Whisper, which perform well under noisy conditions, and the use of Federated Learning to ensure sensitive data remains secure and decentralized, aligning operational needs with regulatory requirements .

Wav2Vec2.0 outperforms Whisper particularly in scenarios involving the GRACE clean corpus for languages such as English, French, and Spanish. This success is attributed to the larger amount of fine-tuning data available for these languages in Wav2Vec2.0's case, which results in a reduction of WER compared to Whisper, despite Whisper's general superiority in diverse and challenging conditions .

The performance of ASR systems like Whisper and Wav2Vec2.0 varies significantly across different languages and datasets. On average, Whisper consistently outperforms Wav2Vec2.0, achieving WERs between 11.5% to 24.9% compared to Wav2Vec2.0's 13.3% to 34.8%. The disparity is more pronounced in low-resource languages, indicating the models' varying adaptability to language-specific challenges. Notably, both models achieve state-of-the-art results for certain languages in the Common Voice, MTEDx, and MediaSpeech corpora .

You might also like