Forensic Speech Recognition in Child Abuse
Forensic Speech Recognition in Child Abuse
v1
Disclaimer/Publisher’s Note: The statements, opinions, and data contained in all publications are solely those of the individual author(s) and
contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting
from any ideas, methods, instructions, or products referred to in the content.
Article
Novel Speech Recognition Systems Applied to Forensics within
Child Exploitation: Wav2vec2.0 vs. Whisper
Juan Camilo Vasquez-Correa1 , Aitor Alvarez-Muniain1
1 Fundación Vicomtech, Basque Research and Technology Alliance (BRTA), Mikeletegi 57, 20009 Donostia-San
Sebastián (Spain)
* Correspondence: {jcvasquez,aalvarez}@[Link] (J.C.V., A.A.)
Abstract: The growth in online child exploitation material is a significant challenge for European 1
Law Enforcement Agencies (LEAs). One of the most important sources of such online information 2
corresponds to audio material that needs to be analyzed to find evidence in a timely and practical 3
manner. That is why LEAs require a next-generation AI-powered platform to process audio data 4
from online sources. We propose the use of speech recognition and keyword spotting to transcribe 5
audiovisual data and to detect the presence of keywords related to child abuse. The considered 6
models are based on two of the most accurate neural-based architectures to date: Wav2vec2.0 7
and Whisper. The systems are tested under an extensive set of scenarios in different languages. 8
Additionally, keeping in mind that obtaining data from LEAs is very sensitive, we explore the use of 9
federated learning to have more robust systems for the addressed application, while maintaining 10
the privacy of the data to LEAs. The considered models achieved a word error rate between 11% 11
and 25%, depending on the language. In addition, the systems are able to recognize a set of spotted 12
words with true positives rates between 82% and 98%, depending on the language. Finally, federated 13
learning strategies show that they can maintain and even improve the performance of the systems 14
when compared to centralized trained models. The proposed systems sit the basis for an AI-powered 15
platform for automatic analysis of audio in the context of forensic applications within child abuse. 16
The use of federated learning is also promising for the addressed scenario, where data privacy is an 17
Keywords: Speech Recognition; Keyword Spotting; Child abuse; Federated Learning; Whisper; 19
Wav2vec2.0 20
1. Introduction 21
The growth in online child exploitation and abuse material is a significant challenge 22
for European Law Enforcement Agencies (LEAs). Currently, the revision of online material 23
about child abuse exceeds the capacity of LEAs to respond in a practical and timely manner. 24
One of the most important sources of information that needs to be analyzed to find evidence 25
about child abuse corresponds to audiovisual material from multimedia content. With the 26
aim to safeguard victims, prosecute offenders and limit the spread of online child abuse 27
data from online sources. One of the main goals of the GRACE project1 is to develop robust 29
AI-based technology to equip LEAs with the aforementioned platform. Two of the core 30
keyword spotting (KWS) in order to accurately transcribe audiovisual online material, and 32
to detect the presence of specific keywords about child abuse in the transcriptions. 33
Within this context, ASR technology has been applied in different forensic scenarios. 34
For instance, to collect evidence via the examination of electronic devices [1], or to analyze 35
1 [Link]
2 of 18
multimedia content related to specific threats [2,3]. Nevertheless, the successful implemen- 36
tation of an ASR system in forensics introduces a series of issues to be solved, which are not 37
present in other domains where ASR is applied. For instance, it is common to find audio 38
coming from different sources, which are highly affected by background noise, overlapping 39
speakers, audio reverberation, among other factors. All these aspects affect the quality of 40
the obtained transcription and the capability of the system to detect specific keywords. 41
novel end-to-end architectures [4] that have shown to be accurate enough in those adverse 43
conditions. The core idea of end-to-end models is to directly map the input speech signal 44
to character sequences and therefore greatly simplify training, fine-tuning and inference [5– 45
9]. Two main approaches are distinguished in the literature to train end-to-end ASR 46
systems: fully supervised or self-supervised models. Regarding the first group, NVIDIA 47
proposed Quartznet [10] with the aim to build a competitive but lighter end-to-end ASR 48
residual connections. The model has been trained and tested on the Common Voice 50
corpus, achieving Word Error Rates (WERs) between 7.7% and 12.5%, depending on the 51
language [11]. A Quartznet model also produced WERs of 19.2% and 18.3% in French 52
and Spanish language multimedia data, respectively, from the MediaSpeech corpus [12]. 53
Researchers from NVIDIA recently proposed Citrinet [13] as an evolution of Quartznet. The 54
authors reported a WER of 5.6% on the TEDLIUMv2 corpus. Another architecture that 57
have proven to be accurate in many ASR benchmark scenarios is the Recurrent Neural 58
Network Transducer (RNN-T) [15]. The RNN-T is formed by three main blocks: (1) an 59
encoder network that receives input acoustic frames and produces high-level speech 60
representations, (2) a predictor that acts as a decoder by processing the previous produced 61
token, and (3) a joint network that combines the outputs from the two previous blocks and 62
produces the distribution of the next predicted token or blank symbol. Recent models based 63
Contrary to fully supervised models, recent studies are focused on the use of big 65
acoustic models trained with self-supervised learning methods and a large amount of 66
unlabelled data. Researchers from Meta AI demonstrated the capabilities of this type of 67
results, especially when considering ASR for low-resource languages in the Common Voice 69
corpus [18]. Particularly, the authors in [19] considered a Wav2vec2.0 model combined with 70
their proposed language modeling approach, and achieve state-of-the-art results in the 71
German Common Voice corpus, with a WER of 3.7%. Wav2Vec2.0-based models have also 72
been successfully tested in more adverse acoustic environments such as in multimedia Por- 73
tuguese data from the CORAA database [20]. Due to these reasons, Wav2Vec2.0 has become 74
one of the most considered neural-based models for ASR. Self-supervised approaches like 75
Wav2Vec2.0 are challenging because there is not a predefined lexicon for the input sound 76
units during the pre-training phase. Moreover, sound units have variable length with no 77
explicit segmentation [21]. With the aim to solve such issues, Meta AI released HuBERT 78
of Convolutional and Transfomer networks from Wav2Vec2.0 and HuBERT has achieved 80
state-of-the-art results in many ASR scenarios. With the aim to combine the best features 81
from both type of networks in a single neural block, researchers from Google introduced the 82
Self-supervised audio encoders like Wav2Vec2.0, HuBERT, and Conformers learn high 85
quality audio representations. However, due to its unsupervised pre-training nature, they 86
lack a proper decoding to transform such representations into usable outputs. This is why 87
a fine-tuning stage is always necessary in order to accurately implement models for ASR 88
or audio classification. With the aim to solve the aforementioned issue, researchers from 89
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
3 of 18
trained in a fully supervised manner, using up to 680,000 hours of labelled audio from 91
internet. The model has achieved state-of-the-art WER results in many benchmark datasets 92
There are two main issues that appear when designing ASR solutions for forensic 94
scenarios: The first one is related to find the most appropriate neural architecture from the 95
ones previously described in order to deal with different acoustic environments. The second 96
one is related to data privacy and protection [26]. Generally, obtaining operative data from 97
LEAs for the addressed scenario is not possible. In this context, Federated Learning (FL) 98
has emerged as an alternative to train machine learning models over remote devices such 99
privacy [27–30]. The procedure is as follows: LEAs operative data are stored in on-premise 101
data servers. Then, FL strategies aim to transfer only local model updates to a central 102
server, keeping LEAs data private. The central server aggregates information obtained 103
from multiple clients i.e., LEAs, and updates a central model that is transmitted back to the 104
clients for their consumption. FL has been applied to train robust federated acoustic models 105
for ASR [31–33] and KWS [34]. In [32] the authors proposed a client adaptive federated 106
training to mitigate data heterogeneity when training ASR models. The proposed system 107
achieved a similar WER with respect to the obtained one using a fully centralized training. 108
In [33] the authors proposed a strategy to compensate Non Independent and Identically 109
Distributed (non-IID) data in federated training of ASR systems. The proposed strategy 110
involved random client data sampling, which resulted in a cost-quality trade-off. The 111
optimization of such a trade-off led to obtain ASRs with similar WERs than the obtained by 112
training centralized systems. The authors in [34] demonstrated the capabilities of federated 113
training to obtain robust KWS systems locally trained on edge devices like smartphones, 114
reaching similar accuracies when compared with centralized trained models. 115
According to the reviewed literature, the two main paradigms and solutions for ASR 116
to date include self-supervised models based on Wav2Vec2.0 and fully supervised models 117
such as Whisper. This work considered and compared these two approaches to test their 118
capabilities to perform robust ASR and KWS in a large set of test scenarios. We also 119
evaluated the use of FL in the context where different LEAs can share a common ASR and 120
KWS system keeping the privacy of their data. In summary, the main contributions of this 121
1. We performed an extensive comparison between two of the most accurate neural- 123
based ASR architectures to date: a fine-tuned version of Wav2Vec2.0 and Whisper. 124
The evaluation is performed in many scenarios, but paying special attention to cor- 125
pora coming from multimedia content. The models are tested in data from seven 126
indo-European Languages, including English, Spanish, German, French, Italian, Por- 127
tuguese, and Polish. This evaluation can be useful as well in other domains besides 128
ASR forensics, making our contribution open and viable in other scenarios. 129
2. We created and released an in domain corpus that includes specific keywords of child 130
abuse domain, and a set of accompanying audios where the keywords are present. 131
The included audios are selected from open available corpora used in the literature. 132
The created corpus can be used as a benchmark to test ASRs in non-controlled acoustic 133
conditions. 134
3. The two neural architectures are compared as well in the created corpora within the 135
scope of child abuse forensics. To the best of our knowledge, this is the first study that 136
comprises the use of open ASR solutions and their capabilities to recognize specific 137
4. We validated the use of FL strategies to train ASR systems in the context of forensic 139
applications. The core idea is that different LEAs can share a common model but 140
The rest of the paper is distributed as follows. Section 2 details different technical 142
aspects of Wav2Vec2.0 and Whisper architectures for ASR. Section 3 describes the consid- 143
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
4 of 18
ered corpora to test the ASR systems, and the process to deliver an in domain corpus for 144
KWS in the context of forensics. Section 4 describes the pilot study on the use of FL for the 145
addressed application. Section 5 displays the main results obtained regarding ASR, KWS, 146
and FL. Section 6 discusses the main insights obtained from the results. Finally, Section 7 147
2. Methods 149
We considered two of the most accurate neural-based ASR architectures to date: (1) 150
Wav2vec2.0, which is trained following a self-supervised paradigm, and (2) Whisper, which 151
is trained following a fully supervised strategy. Details about each model are found in the 152
and Transformer layers (see Figure 1). The model encodes raw audio waveforms χ into 156
Transformer network initially quantise the continuous representations, forming a discrete 159
set of outputs q1 , . . . , q T that represent targets in the self-supervised learning objective [17, 160
35]. Those quantised representations are then contextualised using the attention blocks from 161
The feature encoder is formed by seven convolutional blocks with 512 channels, strides 163
of {5, 2, 2, 2, 2, 2, 2} and kernel widths of {10, 3, 3, 3, 3, 2, 2}. The Transformer network is 164
formed by 24 blocks, a model dimension of 1024, an inner dimension of 4096 and a total of 165
Figure 1. Wav2vec2.0 architecture representation. The raw audio signal is mapped to speech
representations that are fed into a Transformer network to output context representations. Figure
based on the one presented in [17].
XLS-R-300M model, which is available via Hugginface2 . The model was pre-trained in 168
a self-supervised manner using 436k hours of unlabelled speech data in 128 languages 169
2 [Link]
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
5 of 18
from the VoxPopuli [36], Multilingual librispeech (MLS) [37], Common Voice [38], BABEL, 170
and VoxLingua107 [39] corpora. The Wav2Vec2-XLS-R-300M is one of the different versions 171
of the Meta AI’s XLS-R multilingual model [40] composed by 300 million parameters. The 172
multilingual pre-trained model was fine-tuned with labelled speech data (see Section 3.1) in 173
seven languages: English, German, French, Spanish, Italian, Portuguese, and Polish. Each 174
model was trained for 50 epochs, with a batch size of 2, 16 gradient accumulation steps, 175
and a learning rate of 5 × 10−5 , which is warmed up during the initial 10% of the training. 176
The trained acoustic representations are decoded using a Connectionist Temporal 177
Classification (CTC) layer with a beam-search decoding strategy (beam-width=256). The 178
CTC decoding include the use of separate 3-gram language models that are trained using 179
large text corpora, and which are included in the decoding with weights of α = 0.5 and 180
β = 1.5. 181
Whisper is a recently introduced ASR system by OpenAI [25]. Contrary to Wav2vec2.0, 183
Whisper is trained in a fully supervised manner, using up to 680k hours of labelled speech 184
data from multiple sources. The model is based on an encoder-decoder Transformer, which 185
is fed by 80-channel log-Mel spectrograms. The encoder is formed by two convolution 186
layers with a kernel size of 3, followed by a sinusoidal positional encoding, and a stacked 187
set of Transformer blocks. The decoder uses the learned positional embeddings and the 188
same number of Transformer blocks from the encoder. Figure 2 illustrates the general 189
Whisper architecture. Different pre-trained models are available with variations in the 190
number of layers and attention heads. We considered the "Whisper-large" model, which 191
consists of 1550 million parameter distributed in 32 layers and 20 attention heads. The 192
3 [Link]
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
6 of 18
TRANS
EN CRIBE 0.0 The quick brown ...
Next-Token
prediction
MLP
MLP
Self Attention
Cross Attention
Self Attention
Cross attention
Transformer
Encoder Blocks
MLP
MLP Transformer
Self Attention Decoder Blocks
Cross Attention
Self Attention
MLP
Sinusoidal
Positional Cross Attention
Encoding
Self Attention
2x Conv 1D + GELU
Learned
Positional
Encoding
TRANS ...
SOT EN 0.0 the quick
CRIBE
º
Log-Mel Spectogram Tokens in Multitask Training Format
Figure 2. Whisper architecture representation. The log Mel-spectrograms are encoded by a Trans-
former network. Encoded representations are transformed into character outputs and no-speech
tokens via the Transformer decoder. Figure based on the one presented in [25].
The model was not fine-tuned in this study, thus the evaluation for all languages 194
was conducted in a zero-shot setting. The decoding was performed using a beam search 195
strategy with 5 beams, an array of temperature weights of [0.2, 0.4, 0.6, 0.8, 1], and a no 196
repeat n-gram size of 3 in order to take advantage of the language modeling head and to 197
3. Materials 199
This section describes a set of open corpora used to benchmark the two considered 200
ASR systems (Section 3.1), followed by the performed process to derive a set of keywords to 201
be spotted by the considered systems (Section 3.2), and the description of a built in domain 202
dataset considered as well to test the considered models (Section 3.3). 203
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
7 of 18
Table 1. List of public speech corpora considered to test the performance of ASR and KWS systems
based on Wav2Vec2.0 and Whisper.
Test Duration
Corpus name Description Languages
(h)
English 173
German 72
Read sentences collected French 38
Common Voice [38] and validated via Spanish 26
crowd-sourcing Italian 23
Portuguese 6
Polish 7
Spoken Wikipedia Volunteer readers of English 42
Corpus (SWC) [41] Wikipedia articles German 36
Speech segments from French 10
Media Speech [12]
YouTube videos Spanish 10
German 2
Audio recordings French 2
Multilingual TEDx [42] and transcripts Spanish 2
from TED talks Italian 2
Portuguese 2
Audio recordings
TEDLIUMv2 [43] English 3
from TED talks
German 14
French 10
Audio recordings Spanish 10
Multilingual librispeech (MLS) [37]
from audiobooks Italian 5
Polish 2
Portuguese 4
German 3
Crowdsourced French 4
Voxforge read Spanish 5
speech Italian 2
Portuguese 1
Audio recordings from
Debating technologies [44] English 1
transcribed public debates
Recordings from the
Polish Parliamentary corpus [45] Polish 1
Polish parliament
Combination of five
CORAA [20] Portuguese 13
corpora in Portuguese
Different public corpora were considered to train/test the ASR and KWS models. 205
Wav2vec2.0 models were fine-tuned using the Common Voice corpus [38] for each con- 206
sidered language. The amount of available labelled data highly varies depending on the 207
language, and include: 1600 hours for English, 777 hours for German, 623 for French, 324 208
for Spanish, 158 for Italian, 63 for Portuguese, and 43 for Polish. These data are freely 209
available via Huggingface4 . The training data for the Spanish model also included 57 hours 210
from the RTVE2018 dataset [46] from the Albayzin 2018 evaluation challenge. 211
The performance of both the fine-tuned Wav2vec2.0 and Whisper-based models was 212
evaluated in a cross-corpora fashion, considering a large set of databases from the literature 213
that are available in the different languages. The list of considered corpora is observed 214
4 [Link]
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
8 of 18
in Table 1. These corpora were selected in order to test the performance of the models in 215
several recording conditions, which can be closer to the realistic scenarios found by LEAs. 216
Notice that due to the sensitive nature of the target application, it is not possible to get 217
access to realistic operative data from LEAs. However, we created an in-domain synthetic 218
dataset using these open source corpora, which is described in Section 3.3. 219
In order to test the capabilities of the ASR models to spot specific keywords within 221
the child abuse domain, we defined a list of keywords to be spotted. The keyword list was 222
obtained from a set of open documents that include: (1) the "Best Practices on Victim support 223
for LEA first responders" deliverable from the GRACE project5 , (2) the 2021 "Barriers to 224
Compensation for Child Victims of Sexual Exploitation" report from ECPAT6 [47], (3) the 225
study from [48], (4) EUROPOL technical reports [49–51], (5) EUROPOL press-releases 226
from 2018 to 2022 using the keyword "child abuse"7 , (6) Wikipedia articles about "child 227
abuse" and "online child abuse", and (7) UNICEF press-releases about "child abuse"8 . 228
All documents were text crawled and pre-processed by performing lemmatisation, and 229
removing stop words, numbers, and date entities. After this process, we obtained a corpus 230
with 55,059 words, whose 6028 are unique. Figure 3 shows the most important keywords 231
4
Corpus Presence (%)
0
ild
t
ab l
ex use
ontion
ma e
pa al
thr t
eo
b
oto
co net
pr on
r
mi e
rn ol
strphy
tra s
a
ua
gir
ren
ea
no
es
lin
um
we
i
o
iva
vid
ter
i
ch
sch
sex
erc
ph
er
ra
ita
int
og
plo
po
Words
Figure 3. Top 20 of the most important keywords related to child abuse, which were used to test the
capability of the ASR system to detect specific terminology within the domain.
Afterwards, we selected the 100 most repeated words from the corpus, which represent 233
the 33% of the information within the whole set of crawled documents. Finally, we excluded 234
12 terms because they were very broad concepts not related with child abuse, leading to a 235
final set of 88 keywords to be spotted. The obtained keyword list (in English) was translated 236
into the remaining six considered languages in order to have a common benchmark for all 237
languages. 238
We considered an additional corpus to test the implemented ASR systems by merging 240
and filtering the data described in Section 3.1. We selected audio samples from all datasets 241
5 [Link]
6 [Link]
7 [Link]
8 [Link]
5D=
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
9 of 18
that contain at least one of the 88 selected keywords. Table 2 shows the data distribution for 242
each language after selection. The table includes the datasets considered for each language 243
where the keywords are found, the number of utterances, and the total audio duration (in 244
hours). 245
Table 2. Data distribution for the GRACE dataset, which combines different corpora into a single one
within the child abuse domain.
The selected audios were processed in order to have also more realistic acoustic 246
conditions than those expected in forensic applications within the considered domain. 247
The process includes: (1) adding background noise with signal to noise ratios (SNR) 248
between 5 and 30 dB (randomly), (2) adding reverberation using room impulses from the 249
VOiCES dataset [52], and (3) randomly applying the ogg-vorbis codec [53] due to it is 250
commonly found in audio material from online sources. The final ASR and KWS evaluation 251
is performed considering the two versions of the corpus: clean and noisy. This corpus is 252
available online9 to be used as a benchmark dataset for speech recognition in different 253
The considered FL pipeline is performed only with English data and includes five 256
nodes that are used for federated training, a dummy node considered to test the evolution 257
of the learning process, and the central server in charge of aggregating the weights received 258
from the five nodes. Figure 4 shows the implemented architecture. Three of the servers 259
were located at Vicomtech premises (Spain), one server was located at Greece, another one 260
in Portugal, and the remaining one in Cyprus. The aim of these connections is to create 261
a real environment for the pilot, in similar conditions to the expected when the model 262
is trained by different LEAs across Europe. In addition, secure communication between 263
clients and the server was established through a VPN connection to ensure that sensitive 264
data (parameters) are safely transmitted and to prevent unauthorised access. Each node 265
contains data from a different dataset: TEDLIUMv2, debating technologies, Librispeech- 266
other, Librispeech-clean, and SWC. This data configuration aims to evaluate the impact of 267
non-IID data distribution, which is more realistic for the addressed forensic application. 268
9 [Link]
GRACE_ASR.zip
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
10 of 18
site-1
TEDLIUM v2
site-2
site-5 Debating
SWC technologies
dum my node
Libr ispeech test clean
Figure 4. Configuration of the FL architecture. Central server with five client nodes (site-{1, 2, · · · , 5})
and a dummy node only used to test the performance of the aggregated model
The FL pilot was performed only with the Wav2Vec2.0 model, and using also the 269
pre-trained Wav2Vec2-XLS-R-300M model. The training hyperparameters were the same 270
for the five clients, and include a batch size of 2, a learning rate of 5 × 10−5 warmed up 271
in the first 10% of the training time, and a gradient accumulation of 16 steps. The local 272
training is performed for 5 epochs. The central server is configured to run for 10 rounds of 273
federated training, and using the Federated averaging (FedAvg) aggregation mechanism 274
to update the central model. The architecture configuration and the training process is 275
5. Results 277
Wav2Vec2.0 and Whisper models were evaluated under the described corpora in 279
Section 3.1. The results of the ASR systems in terms of WER are shown in Table 3. The 280
results included those obtained in the evaluation of the seven languages, and using both the 281
open benchmark corpora and the two versions (clean and noisy) of the synthetic GRACE 282
corpus. 283
10 [Link]
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
11 of 18
Table 3. Results of the ASR models in different languages considering all benchmark datasets. Results
in terms of WER.
Model Common MLS TED- MTEDx SWC Media Voxforge Debates Polish CORAA GRACE GRACE AVG.
Voice LIUMv2 Speech Parl clean noisy
English
Wav2Vec 2.0 16.1 - 17.2 - 20.6 - - 11.7 - - 18.9 32.6 19.5
Whisper 10.0 - 5.4 - 20.6 - - 7.0 - - 24.5 19.8 14.6
German
Wav2Vec 2.0 11.9 12.9 - 36.7 34.5 - 7.5 - - - 20.0 33.8 22.5
Whisper 7.1 6.7 - 21.7 18.3 4.2 - - - 15.5 22.5 13.7
French
Wav2Vec 2.0 16.7 17.0 - 25.3 - 29.1 16.7 - - - 26.5 56.3 26.8
Whisper 21.7 8.0 - 23.3 - 35.8 14.6 - - - 36.8 34.1 24.9
Spanish
Wav2Vec 2.0 4.7 7.2 - 12.9 - 14.5 6.3 - - - 12.6 33.3 13.1
Whisper 6.2 5.3 - 9.4 - 15.8 4.2 - - - 19.6 18.8 11.3
Italian
Wav2Vec 2.0 12.8 21.1 - 22.2 - - 14.3 - - - 18.0 46.3 22.5
Whisper 7.9 13.6 - 11.6 - - 10.5 - - - 14.1 20.2 13.0
Portuguese
Wav2Vec 2.0 12.9 20.1 - 33.8 - - 17.8 - - 48.5 42.7 68.1 34.8
Whisper 5.4 8.8 - 13.1 - - 11.2 - - 21.7 22.1 42.3 17.8
Polish
Wav2Vec 2.0 11.5 12.7 - - - - - - 32.1 - - - 18.8
Whisper 8.9 6.0 - - - - - - 32.5 - - - 15.8
On average, the WER for each language using Whisper ranges from 11.3% (in Spanish) 284
to 24.9% (in French). The results using Wav2Vec2.0 range from 13.1% (in Spanish) to 34.8% 285
(in Portuguese). In general, Whisper produces less errors than Wav2Vec2.0 (see Figure 5 286
left). The difference between both models is statistically significant according to a Mann 287
under the most affected acoustic conditions, such as in the GRACE noisy, TEDLIUMv2, 289
Debates, and CORAA corpora. However, there are some scenarios where Wav2Vec2.0 290
outperformed Whisper and which should be considered with special attention, such as the 291
The results obtained were compared to those found in the literature for the multilingual 293
corpora: Common Voice, MLS, MTEDx, and MediaSpeech. The comparison is shown in 294
Table 4. The Wav2Vec2.0-based model outperformed results in the Spanish versions of 295
Common Voice and MediaSpeech corpora, with WERs of 4.3% and 14.5%, respectively 296
with respect to to the results reported in [18] for Common Voice (WER=6.2%) and in [12] 297
for MediaSpeech (WER=18.3%). We also reported state-of-the-art results for the Spanish, 298
Portuguese, Italian, and German versions of the MTEDx corpus (WERs of 9.4%, 12%, 299
11.6%, and 21.7%, respectively) with respect to the WERs of 16.2%, 20.2%, 16.4%, and 300
42.3% reported in [42]. Whisper model also achieved state-of-the-art results in the CORAA 301
corpus (WER=21.7%) with respect to the results reported in [54] (WER=21.9%), and in 302
the TEDLIUMv2 corpus (WER=5.4%) compared to [13] (WER=5.6%). Regarding MLS, the 303
state-of-the-art results are still from [55]. However, notice that the results reported here 304
correspond to cross-corpus tests, while the experiments performed in [55] correspond to 305
Wav2Vec2.0 models trained and tested using MLS, thus making the models adapted just 306
12 of 18
Table 4. WER comparison between the results reported and those coming from the state-of-the-art for
Common Voice, MLS, and MTEDx corpora. Best results for each corpus and language are highlighted
in bold
The text transcriptions from Wav2Vec2.0 and Whisper were post-processed in order to 309
find the presence of the defined keywords to be spotted. The process involved transforming 310
the inflectional form of each word in order to detect all possible variations of the word 312
within the transcription. The lemmatization process is performed using the set of large 313
open dictionaries available in Spacy11 . The results obtained for KWS in each corpus are 314
shown in Table 5. The results are presented in terms of the true positive rate (TPR). This is 315
a common metric used in this type of applications where it is more important to avoid false 316
On average, the TPRs are higher using Whisper, and the results per language using 318
Whisper range from 81.5% (for Polish) to 98.4% (for Italian). Results using Wav2Vec2.0 319
range from 82.9% (for Portuguese) to 94.9% (for Spanish). Similar to the ASR results, the 320
difference between Whisper and Wav2Vec2.0 is larger when considering speech signals 321
in non-controlled acoustic conditions, like the ones from the GRACE noisy corpus, where 322
we particularly guarantee the presence of the spotted keywords in every utterance. High 323
differences were also observed int the CORAA corpus, in Common Voice, and in the 324
German SWC. The difference between the results obtained using Wav2Vec2.0 and Whisper 325
is also statistically significant (see Figure 5 right) according to a Mann-Whitney test with 326
11 [Link]
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
13 of 18
Table 5. Results of KWS in different languages considering all benchmark datasets. Results in terms
of TPR (%).
Model Common MLS TED- MTEDx SWC Media Voxforge Debates Polish CORAA GRACE GRACE AVG.
Voice LIUMv2 Speech Parl clean noisy
English
Wav2Vec 2.0 93.3 - 95.4 - 92.5 - - 96.6 - - 94.6 79.5 92.0
Whisper 96.8 - 97.4 - 94.1 - - 97.7 - - 91.8 93.6 95.2
German
Wav2Vec 2.0 91.3 96.9 - 93.5 80.4 - 99.8 - - - 94.3 79.7 90.8
Whisper 97.8 98.8 - 97.7 97.6 99.8 - - - 97.2 90.6 97.1
French
Wav2Vec 2.0 90.6 90.1 - 94.5 - 82.8 90.7 - - - 90.5 60.1 85.6
Whisper 94.6 98.0 - 93.9 - 84.9 94.2 - - - 88.0 81.3 90.7
Spanish
Wav2Vec 2.0 96.7 98.1 - 98.1 - 96.3 100.0 - - - 97.0 78.1 94.9
Whisper 98.2 99.8 - 98.6 - 94.4 99.8 - - - 92.0 94.2 96.7
Italian
Wav2Vec 2.0 90.7 97.2 - 95.5 - - 98.9 - - - 96.1 80.9 93.2
Whisper 97.8 99.9 - 97.3 - - 99.8 - - - 98.7 96.9 98.4
Portuguese
Wav2Vec 2.0 93.1 93.4 - 94.1 - - 99.1 - - 74.3 76.8 49.5 82.9
Whisper 96.6 97.5 - 99.4 - - 100.0 - - 88.1 88.3 81.3 93.0
Polish
Wav2Vec 2.0 93.9 96.9 - - - - - - 83.3 - - - 91.4
Whisper 95.4 98.7 - - - - - - 50.3 - - - 81.5
0.04 0.125
Normalized count
Normalized count
0.100
0.03
0.075
0.02
0.050
0.01 0.025
0.00 0.000
0 10 20 30 40 50 60 70 80 30 40 50 60 70 80 90 100
WER (%) TPR (%)
Figure 5. Comparison between the results obtained using Wav2Vec2.0 and Whisper for ASR (left)
and KWS (right).
The FL experiment involved training the Wav2Vec2.0 system using 5 separate real 329
servers for training, and one additional node (dummy) used only to test the final model. 330
Each node contained data from a different dataset (only in English) in order to evaluate 331
the contribution from each corpus into the global aggregated model. The aim was also to 332
cover non-IID conditions, which have shown to be one of the most important drawbacks 333
when training models in an FL approach. The results are shown in Table 6. The results 334
using the FL training are compared to those obtained training the system in a complete 335
centralized manner. Similar WERs were obtained in each node comparing the federated 336
vs. centralized training. The main difference is that when considering FL models there is 337
only one aggregated model which covers the results of the 5 nodes, instead of having 5 338
different models for the case of the centralized approach. This fact highly reduces the time 339
considered to train the system, and most important, it is possible to take advantage of data 340
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
14 of 18
from different data centers to train a more robust and general model without the need of 341
Table 6. Results of the FL pilot comparing WERs from Wav2Vec2.0 models trained in a federated or
centralized way.
6. Discussion 343
The evaluation of Wav2Vec2.0 and Whisper-based ASR systems is performed under a 344
large set of different scenarios, including one specifically designed for forensic applications 345
within child domain abuse. On average, Whisper is more accurate than the Wav2Vec2.0- 346
based system. Whisper achieved WERs ranging from 11.5% to 24.9%, depending on the 347
language, compared with Wav2Vec2.0 WERs that range between 13.3% to 34.8%. The 348
difference between the two models is even larger when considering languages trained with 349
lower resources, such as Portuguese or Italian. Although these differences, Wav2vec2.0 350
is competitive with Whisper when the number of hours for fine-tuning is large, e.g, in 351
Results using the GRACE dataset show relatively similar WERs between Wav2Vec2.0 353
and Whisper when considering the clean version of the corpus, with an average WER of 354
22.1% for Whisper and of 23.2% for Wav2Vec2.0. However, the difference between the 355
two models greatly increases when considering the noisy version of the corpus, with an 356
average WER of 26.3% for Whisper and of 45.1% for Wav2Vec2.0. This is a great indicator 357
about the capability of Whisper to perform accurate transcriptions under non-controlled 358
and noisy acoustic conditions, by keeping similar WERs in the two versions of the GRACE 359
corpus. Despite the differences between the two types of models, there are some surprising 360
results where Wav2Vec2.0 outperforms Whisper, and which should be considered with 361
special attention. For instance, when evaluating the GRACE clean corpus in languages such 362
as English, French, and Spanish. The models for these three languages were fine-tuned 363
with more data, which likely explains the WER reduction in Wav2Vec2.0 with respect to 364
Whisper. 365
The performed evaluations of our systems achieved state-of-the art results in several 366
of the considered benchmark corpora. We reported state-of-the-art results for some of the 367
languages in the Common Voice corpus. State-of-the-art results were also achieved for 368
almost all languages in the MTEDx and MediaSpeech corpora. These results are good 369
indicators about the capabilities of the considered systems to accurately recognize speech 370
under more natural and spontaneous scenarios, closer to the expected in forensic domains. 371
The KWS evaluation indicated that both Wav2Vec2.0 and Whisper were accurate 372
enough to recognize the considered child abuse-related keywords in the seven languages. 373
TPRs obtained for Wav2Vec2.0 range from 82.9% to 94.9%, depending on the language. 374
Results using Whisper range from 80.3% to 98.2%. The particular evaluation of KWS in the 375
GRACE dataset also shows that both models are equally accurate to recognize the selected 376
keywords under controlled acoustic conditions. On the contrary, when considering the 377
noisy version of the corpus, the results for Wav2Vec2.0 are reduced by 20% while the results 378
for Whisper are only reduced by 3%. This fact again indicates the capability of Whisper to 379
The last covered experiment involved a pilot study on the use of FL to train ASR 381
systems. The results indicated that an ASR trained in a federated way maintains and 382
in some cases outperforms the performance of individual ASRs trained in a centralized 383
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
15 of 18
manner by each LEA. In addition to the performance, the most important aspect of FL is 384
that the ASR training does not involve any data sharing among LEAs, since only updates 385
of the network parameters are transferred to a central server in charge of aggregating the 386
model. These results are indicators about the potential use of FL to obtain a joint (and 387
potentially richer) model combining sources of data that could not be otherwise combined. 388
Although the benefits of using FL, it is important to consider external factors that may 389
degrade the performance and reliability of the system. For instance, there is evidence about 390
FL attacks that are able to retrieve speaker information from the transferred weights [60] 391
or data poisoning attacks inside LEAs server. Different strategies can be considered to 392
mitigate this this type of attacks such as the use of differential privacy algorithms [61] or 393
7. Conclusions 395
This paper proposed the use of speech recognition and keyword spotting technologies 396
to be applied in forensic scenarios, particularly in child exploitation domains. The aim is 397
to provide LEAs with technology to detect the presence of offensive online audiovisual 398
material related to child abuse. State-of-the art ASR systems based on Wav2Vec2.0 and 399
Whisper were considered for the addressed application. The performance of both models 400
was tested on a large set of open benchmark corpora from the literature. Therefore, the 401
results obtained can be extended to other ASR domains. We additionally created an in- 402
domain corpus using different open source datasets from the research community. The aim 403
was to test the models in more realistic and operative conditions. 404
The ASR and KWS models were evaluated in corpora from seven Indo-European 405
languages, including English, German, French, Spanish, Italian, Portuguese, and Polish. 406
We obtained overall WERs ranging from 11.3% to 24.9%, depending on the language. The 407
performance of the KWS model for the different languages ranged from 81.5% to 98.4%. 408
The most accurate results were obtained from models trained with more data, such as 409
English or German. The comparison between Wav2Vec2.0 and Whisper models indicated 410
that the second one was the most accurate system in the majority of cases, especially when 411
We also proposed a strategy for using FL to train robust ASR systems in the context 413
of the addressed application. This is a suitable approach considering that collecting op- 414
erational data from LEAs is not possible. FL approaches allow LEAs to build a common 415
technological platform without the need to share their operational data. The results of the 416
FL pilot indicated that similar WERs were achieved when comparing the model trained in 417
a federated way to individual models trained in a centralized manner, even considering 418
non-IID conditions, which has been shown to be one of the main drawbacks in FL. 419
For future work, the considered approaches can be extended to other forensic applica- 420
tions where there is a need to monitor audiovisual material from online sources. In addition, 421
the considered technology can be combined with other speech processing methods, such as 422
speaker and language identification, age and gender recognition, and speaker diarization. 423
The ultimate goal is to provide LEAs with accurate tools to monitor audio from online 424
Author Contributions: "Conceptualization, J.C.V., A.A., and .; methodology, J.C.V. and A.A; software, 426
J.C.V.; validation, J.C.V.; formal analysis, J.C.V. and A.A; investigation, J.C.V, A.A., and .; resources, .; 427
data curation, J.C.V.; writing—original draft preparation, J.C.V.; writing—review and editing, J.C.V., 428
A.A., and .; visualization, J.C.V. All authors have read and agreed to the published version of the 429
manuscript.” 430
Funding: This project has received funding from the European Union’s Horizon 2020 research and 431
Institutional Review Board Statement: The study was conducted in accordance with the Declaration 433
of Helsinki, and approved by the Institutional Review Board of the GRACE consortium (protocol 434
16 of 18
Data Availability Statement: All data considered in this study come from open repositories under 436
Abbreviations 439
References 441
1. Negrão, M.; Domingues, P. SpeechToText: An open-source software for automatic detection and transcription of voice recordings 442
in digital forensics. Forensic Science International: Digital Investigation 2021, 38, 301223. 443
2. Alghowinem, S. A safer youtube kids: An extra layer of content filtering using automated multimodal analysis. In Proceedings 444
of the Proceedings of SAI Intelligent Systems Conference. Springer, 2018, pp. 294–308. 445
3. Mariconti, E.; Suarez-Tangil, G.; Blackburn, J.; De Cristofaro, E.; Kourtellis, N.; Leontiadis, I.; Serrano, J.L.; Stringhini, G. " You 446
Know What to Do" Proactive Detection of YouTube Videos Targeted by Coordinated Hate Attacks. Proceedings of the ACM on 447
4. Amodei, D.; Ananthanarayanan, S.; Anubhai, R.; Bai, J.; Battenberg, E.; Case, C.; Casper, J.; Catanzaro, B.; Cheng, Q.; Chen, G.; 449
et al. Deep speech 2: End-to-end speech recognition in english and mandarin. In Proceedings of the International conference on 450
5. Graves, A.; Jaitly, N. Towards end-to-end speech recognition with recurrent neural networks. In Proceedings of the International 452
6. Chan, W.; Jaitly, N.; Le, Q.; Vinyals, O. Listen, attend and spell: A neural network for large vocabulary conversational speech 454
7. Chorowski, J.K.; Bahdanau, D.; Serdyuk, D.; Cho, K.; Bengio, Y. Attention-based models for speech recognition. Advances in 456
8. Lu, L.; Zhang, X.; Renais, S. On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech 458
9. Yao, Z.; Wu, D.; Wang, X.; Zhang, B.; Yu, F.; Yang, C.; Peng, Z.; Chen, X.; Xie, L.; Lei, X. Wenet: Production oriented streaming and 460
non-streaming end-to-end speech recognition toolkit. arXiv preprint arXiv:2102.01547 2021. 461
10. Kriman, S.; Beliaev, S.; Ginsburg, B.; Huang, J.; Kuchaiev, O.; Lavrukhin, V.; Leary, R.; Li, J.; Zhang, Y. Quartznet: Deep automatic 462
speech recognition with 1d time-channel separable convolutions. In Proceedings of the ICASSP. IEEE, 2020, pp. 6124–6128. 463
11. Bermuth, D.; Poeppel, A.; Reif, W. Scribosermo: Fast Speech-to-Text models for German and other Languages. arXiv preprint 464
12. Kolobov, R.; Okhapkina, O.; Omelchishina, O.; Platunov, A.; Bedyakin, R.; Moshkin, V.; Menshikov, D.; Mikhaylovskiy, N. 466
Mediaspeech: Multilanguage asr benchmark and dataset. arXiv preprint arXiv:2103.16193 2021. 467
13. Majumdar, S.; Balam, J.; Hrinchuk, O.; Lavrukhin, V.; Noroozi, V.; Ginsburg, B. Citrinet: Closing the gap between non- 468
autoregressive and autoregressive end-to-end models for automatic speech recognition. arXiv preprint arXiv:2104.01721 2021. 469
14. Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the Proceedings of the IEEE conference on computer 470
15. Graves, A. Sequence transduction with recurrent neural networks. arXiv preprint arXiv:1211.3711 2012. 472
16. Zhou, W.; Zheng, Z.; Schlüter, R.; Ney, H. On language model integration for rnn transducer based speech recognition. In 473
17. Baevski, A.; et al. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information 475
18. Pham, N.Q.; Waibel, A.; Niehues, J. Adaptive multilingual speech recognition with pretrained models. In Proceedings of the 477
19. Krabbenhöft, H.N.; Barth, E. TEVR: Improving Speech Recognition by Token Entropy Variance Reduction. arXiv preprint 479
20. Junior, A.C.; Casanova, E.; Soares, A.; de Oliveira, F.S.; Oliveira, L.; Junior, R.C.F.; da Silva, D.P.P.; Fayet, F.G.; Carlotto, B.B.; Gris, 481
L.R.S.; et al. CORAA: a large corpus of spontaneous and prepared speech manually validated for speech recognition in Brazilian 482
17 of 18
21. Hsu, W.N.; Tsai, Y.H.H.; Bolte, B.; Salakhutdinov, R.; Mohamed, A. HuBERT: How much can a bad teacher benefit ASR 484
22. Hsu, W.N.; Bolte, B.; Tsai, Y.H.H.; Lakhotia, K.; Salakhutdinov, R.; Mohamed, A. Hubert: Self-supervised speech representation 486
learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing 2021, 29, 3451– 487
3460. 488
23. Gulati, A.; Qin, J.; Chiu, C.C.; Parmar, N.; Zhang, Y.; Yu, J.; Han, W.; Wang, S.; Zhang, Z.; Wu, Y.; et al. Conformer: Convolution- 489
augmented Transformer for Speech Recognition. In Proceedings of the Proc. Interspeech 2020, 2020, pp. 5036–5040. https: 490
//[Link]/10.21437/Interspeech.2020-3015. 491
24. Guo, P.; Boyer, F.; Chang, X.; Hayashi, T.; Higuchi, Y.; Inaguma, H.; Kamo, N.; Li, C.; Garcia-Romero, D.; Shi, J.; et al. Recent 492
developments on espnet toolkit boosted by conformer. In Proceedings of the ICASSP. IEEE, 2021, pp. 5874–5878. 493
25. Radford, A.; Kim, J.W.; Xu, T.; Brockman, G.; McLeavey, C.; Sutskever, I. Robust speech recognition via large-scale weak 494
26. Voigt, P.; Von dem Bussche, A. The eu general data protection regulation (gdpr). A Practical Guide, 1st Ed., Cham: Springer 496
27. Konečnỳ, J.; McMahan, H.B.; Yu, F.X.; Richtárik, P.; Suresh, A.T.; Bacon, D. Federated learning: Strategies for improving 498
28. Yang, Q.; Liu, Y.; Cheng, Y.; Kang, Y.; Chen, T.; Yu, H. Federated learning. Synthesis Lectures on Artificial Intelligence and Machine 500
29. Li, L.; Fan, Y.; Tse, M.; Lin, K.Y. A review of applications in federated learning. Computers & Industrial Engineering 2020, 502
30. Li, T.; Sahu, A.K.; Talwalkar, A.; Smith, V. Federated learning: Challenges, methods, and future directions. IEEE Signal Processing 504
31. Dimitriadis, D.; Kumatani, K.; Gmyr, R.; Gaur, Y.; Eskimez, S.E. A Federated Approach in Training Acoustic Models. In 506
32. Cui, X.; Lu, S.; Kingsbury, B. Federated acoustic modeling for automatic speech recognition. In Proceedings of the ICASSP. IEEE, 508
33. Guliani, D.; Beaufays, F.; Motta, G. Training speech recognition models with federated learning: A quality/cost framework. In 510
34. Hard, A.; Partridge, K.; Nguyen, C.; Subrahmanya, N.; Shah, A.; Zhu, P.; Moreno, I.L.; Mathews, R. Training Keyword 512
Spotting Models on Non-IID Data with Federated Learning. In Proceedings of the INTERSPEECH, 2020, pp. 4343–4347. 513
[Link] 514
35. Conneau, A.; Baevski, A.; Collobert, R.; Mohamed, A.; Auli, M. Unsupervised cross-lingual representation learning for speech 515
36. Wang, C.; Riviere, M.; Lee, A.; Wu, A.; Talnikar, C.; Haziza, D.; Williamson, M.; Pino, J.; Dupoux, E. VoxPopuli: A Large-Scale 517
Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation. In Proceedings of 518
the Annual Meeting of the Association for Computational Linguistics and International Joint Conference on Natural Language 519
37. Pratap, V.; Xu, Q.; Sriram, A.; Synnaeve, G.; Collobert, R. MLS: A Large-Scale Multilingual Dataset for Speech Research. In 521
38. Ardila, R.; Branson, M.; Davis, K.; Henretty, M.; Kohler, M.; Meyer, J.; Morais, R.; Saunders, L.; Tyers, F.M.; Weber, G. Common 523
Voice: A Massively-Multilingual Speech Corpus. In Proceedings of the Proceedings of the 12th Conference on Language 524
39. Valk, J.; Alumäe, T. VoxLingua107: a dataset for spoken language recognition. In Proceedings of the IEEE Spoken Language 526
40. Babu, A.; Wang, C.; Tjandra, A.; Lakhotia, K.; Xu, Q.; Goyal, N.; Singh, K.; von Platen, P.; Saraf, Y.; Pino, J.; et al. XLS-R: Self- 528
supervised Cross-lingual Speech Representation Learning at Scale. In Proceedings of the INTERSPEECH, 2022, pp. 2278–2282. 529
[Link] 530
41. Baumann, T.; Köhn, A.; Hennig, F. The Spoken Wikipedia Corpus collection: Harvesting, alignment and an application to 531
42. Salesky, E.; Wiesner, M.; Bremerman, J.; Cattoni, R.; Negri, M.; Turchi, M.; Oard, D.W.; Post, M. The Multilingual TEDx Corpus 533
for Speech Recognition and Translation. In Proceedings of the INTERSPEECH, 2021, pp. 3655–3659. [Link] 534
/Interspeech.2021-11. 535
43. Rousseau, A.; Deléglise, P.; Esteve, Y.; et al. Enhancing the TED-LIUM corpus with selected data for language modeling and more 536
44. Mirkin, S.; Jacovi, M.; Lavee, T.; Kuo, H.K.; Thomas, S.; Sager, L.; Kotlerman, L.; Venezian, E.; Slonim, N. A Recorded Debating 538
45. Ogrodniczuk, M. Polish parliamentary corpus. In Proceedings of the LREC, 2018, pp. 15–19. 540
46. Lleida, E.; Ortega, A.; Miguel, A.; Bazán-Gil, V.; Pérez, C.; Gómez, M.; De Prada, A. Albayzin 2018 evaluation: the iberspeech-rtve 541
challenge on speech technologies for spanish broadcast media. Applied sciences 2019, 9, 5412. 542
Preprints ([Link]) | NOT PEER-REVIEWED | Posted: 22 December 2022 doi:10.20944/preprints202212.0426.v1
18 of 18
47. ECPAT. Barriers to Compensation for Child Victims of Sexual Exploitation A discussion paper based on a comparative legal 543
48. Richards, K. Misperceptions about child sex offenders. Trends and issues in crime and criminal justice 2011, pp. 1–8. 545
49. EUROPOL. Online sexual coercion and extortion as a form of crime affecting children. European Union Agency for Law Enforcement 546
50. EUROPOL. Internet Organised Crime Threat Assessment. European Union Agency for Law Enforcement Cooperation 2019. 548
51. EUROPOL. Exploting Isolation: Offenders and victims of online child sexual abuse during the COVID-19 pandemic. European 549
52. Richey, C.; Barrios, M.A.; Armstrong, Z.; Bartels, C.; Franco, H.; Graciarena, M.; Lawson, A.; Nandwana, M.K.; Stauffer, A.; van 551
Hout, J.; et al. Voices Obscured in Complex Environmental Settings (VOiCES) Corpus. In Proceedings of the INTERSPEECH, 552
53. Moffitt, J. Ogg Vorbis—open, free audio—set your media free. Linux journal 2001, 2001, 9–es. 554
54. Marcacini, R.M.; Candido Junior, A.; Casanova, E. Overview of the Automatic Speech Recognition for Spontaneous and Prepared 555
Speech & Speech Emotion Recognition in Portuguese (SE&R) Shared-tasks at PROPOR 2022. In Proceedings of the PROPOR, 556
2022. 557
55. Bai, J.; Li, B.; Zhang, Y.; Bapna, A.; Siddhartha, N.; Sim, K.C.; Sainath, T.N. Joint unsupervised and supervised training for 558
multilingual asr. In Proceedings of the ICASSP. IEEE, 2022, pp. 6402–6406. 559
56. Zheng, H.; Peng, W.; Ou, Z.; Zhang, J. Advancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces and 560
57. Stefanel Gris, L.R.; Casanova, E.; Oliveira, F.S.d.; Silva Soares, A.d.; Candido Junior, A. Brazilian Portuguese Speech Recognition 562
Using Wav2vec 2.0. In Proceedings of the International Conference on Computational Processing of the Portuguese Language. 563
58. Keshet, J.; Grangier, D.; Bengio, S. Discriminative keyword spotting. Speech Communication 2009, 51, 317–329. 565
59. Lengerich, C.; Hannun, A. An end-to-end architecture for keyword spotting and voice activity detection. arXiv preprint 566
60. Tomashenko, N.; Mdhaffar, S.; Tommasi, M.; Estève, Y.; Bonastre, J.F. Privacy attacks for automatic speech recognition acoustic 568
models in a federated learning framework. In Proceedings of the ICASSP. IEEE, 2022, pp. 6972–6976. 569
61. Geyer, R.C.; Klein, T.; Nabi, M. Differentially private federated learning: A client level perspective. arXiv preprint arXiv:1712.07557 570
2017. 571
Federated Learning has the potential to significantly impact the collaborative development of ASR technology among law enforcement agencies by enabling shared model improvements without compromising data security. It facilitates the combination of models trained on diverse datasets, leading to richer, more robust systems. This collaboration can yield models that are better suited to various dialects and conditions met by different agencies while maintaining strict data privacy, thus enhancing ASR's applicability in widespread forensic operations .
Both Whisper and Wav2Vec2.0 are effective in recognizing child abuse-related keywords across different languages, achieving significant True Positive Rates (TPRs). Wav2Vec2.0's TPRs range from 82.9% to 94.9%, while Whisper's range is 80.3% to 98.2%. These results underscore Whisper's slightly superior performance in handling various languages, likely due to its large training dataset and robust handling of noise .
Federated Learning (FL) addresses privacy concerns by decentralizing the training process, allowing models to be trained on local devices without the need to share sensitive data. Instead of sending raw data, only updates of the network parameters are transferred to a central server, which aggregates these updates to improve a global model. This method keeps the data on-premise, ensuring that sensitive information relating to law enforcement activities is not transmitted elsewhere, thereby maintaining privacy and data security .
Federated Learning (FL) enhances the robustness of ASR models by allowing them to be trained on diverse datasets from multiple sources without centralizing the actual data. This approach helps overcome data heterogeneity issues, as insights and patterns from various local datasets are aggregated into a central model, providing a more generalized model without compromising data privacy. FL has been shown to maintain, and in some cases exceed, the performance of models trained centrally .
Whisper is considered more suitable for forensic applications primarily due to its superior performance under noisy and non-controlled environments, which are common in forensic settings. The model achieves lower WERs in these challenging acoustic environments, maintaining a consistent level of accuracy where Wav2Vec2.0 performance falls considerably .
The Whisper model demonstrates a significant advantage in noisy acoustic conditions compared to Wav2Vec2.0. When tested on the GRACE dataset, the average WER for Whisper in noisy conditions was 26.3% compared to 45.1% for Wav2Vec2.0. This showcases Whisper's strong capability to maintain accuracy in non-controlled environments, reducing the impact of noise substantially less than Wav2Vec2.0 does .
Data heterogeneity in federated training poses a challenge as it can lead to inefficiencies and degraded model performance if the data distribution varies significantly across clients. To address this, researchers have implemented client adaptive federated training strategies and random client data sampling methods. These approaches aim to balance the quality-cost trade-off and compensate for non-IID data, thereby achieving comparable WERs to fully centralized systems .
When designing ASR solutions for forensic scenarios, the main considerations include selecting neural architectures that can handle various acoustic environments and ensuring data privacy and protection. These considerations lead to the preference for models like Whisper, which perform well under noisy conditions, and the use of Federated Learning to ensure sensitive data remains secure and decentralized, aligning operational needs with regulatory requirements .
Wav2Vec2.0 outperforms Whisper particularly in scenarios involving the GRACE clean corpus for languages such as English, French, and Spanish. This success is attributed to the larger amount of fine-tuning data available for these languages in Wav2Vec2.0's case, which results in a reduction of WER compared to Whisper, despite Whisper's general superiority in diverse and challenging conditions .
The performance of ASR systems like Whisper and Wav2Vec2.0 varies significantly across different languages and datasets. On average, Whisper consistently outperforms Wav2Vec2.0, achieving WERs between 11.5% to 24.9% compared to Wav2Vec2.0's 13.3% to 34.8%. The disparity is more pronounced in low-resource languages, indicating the models' varying adaptability to language-specific challenges. Notably, both models achieve state-of-the-art results for certain languages in the Common Voice, MTEDx, and MediaSpeech corpora .