0% found this document useful (0 votes)
10 views15 pages

Child Speech Synthesis TTS Pipeline Evaluation

speech therapy

Uploaded by

chn22csd229
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views15 pages

Child Speech Synthesis TTS Pipeline Evaluation

speech therapy

Uploaded by

chn22csd229
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Received April 2, 2022, accepted April 20, 2022, date of publication April 28, 2022, date of current version

May 6, 2022.
Digital Object Identifier 10.1109/ACCESS.2022.3170836

A Text-to-Speech Pipeline, Evaluation


Methodology, and Initial Fine-Tuning Results for
Child Speech Synthesis
RISHABH JAIN 1 , (Graduate Student Member, IEEE), MARIAM YAHAYAH YIWERE1 ,
DAN BIGIOI 1 , (Graduate Student Member, IEEE), PETER CORCORAN 1 , (Fellow, IEEE),
AND HORIA CUCU 2 , (Member, IEEE)
1 School of Electrical and Electronics Engineering, National University of Ireland Galway, Galway, H91 TK33 Ireland
2 Speech and Dialogue Research Laboratory, University Politehnica of Bucharest, RO-060042 Bucharest, Romania
Corresponding author: Rishabh Jain ([Link]@[Link])
This work was supported by the Data-Center Audio/Visual Intelligence on-Device (DAVID) Project (2020–2023) funded by the Disruptive
Technologies Innovation Fund (DTIF), Established under Project Ireland 2040 through the Department of Enterprise, Trade, and
Employment with Administrative Support from Enterprise Ireland.

ABSTRACT Speech synthesis has come a long way as current text-to-speech (TTS) models can now generate
natural human-sounding speech. However, most of the TTS research focuses on using adult speech data
and there has been very limited work done on child speech synthesis. This study developed and validated a
training pipeline for fine-tuning state-of-the-art (SOTA) neural TTS models using child speech datasets. This
approach adopts a multi-speaker TTS retuning workflow to provide a transfer-learning pipeline. A publicly
available child speech dataset was cleaned to provide a smaller subset of approximately 19 hours, which
formed the basis of our fine-tuning experiments. Both subjective and objective evaluations were performed
using a pretrained MOSNet for objective evaluation and a novel subjective framework for mean opinion
score (MOS) evaluations. Subjective evaluations achieved the MOS of 3.95 for speech intelligibility, 3.89 for
voice naturalness, and 3.96 for voice consistency. Objective evaluation using a pretrained MOSNet showed a
strong correlation between real and synthetic child voices. Speaker similarity was also verified by calculating
the cosine similarity between the embeddings of utterances. An automatic speech recognition (ASR) model
is also used to provide a word error rate (WER) comparison between the real and synthetic child voices. The
final trained TTS model was able to synthesize child-like speech from reference audio samples as short as
5 seconds.

INDEX TERMS Text-to-speech, child speech synthesis, tacotron, multi-speaker TTS, alternative WaveRNN,
MOSNet, subjective MOS.

I. INTRODUCTION voice services, TTS models are also important, and the most
The bulk of recent research into human speech has focused on advanced models can incorporate emotional and prosodic
neural network techniques to improve speech understanding elements into the generated speech output.
and recognition or to provide simplified, high-quality text- More recent research into low-resource languages and
to-speech (TTS) models that can directly convert written other low-resource aspects of human speech, such as accented
text into natural speech. The most highly developed domain and prosody-aligned speech has started to see improvements
for such research has a focus on spoken English and is for both ASR and TTS [1]. Another aspect of human speech
based on native-speaker adult voice data samples. Automated of growing importance is that of child speech. Child speech
speech recognition (ASR) is a core element of modern con- differs significantly from those of adult speech, falling into a
sumer technology user interfaces employed in smart-speaker narrow range of variation and with higher pitch levels. Fur-
and voice command interfaces. For interactive chatbot and thermore, children’s speech patterns are more inarticulate and
can vary widely in terms of volume, pacing, and emotional
expressivity. These challenges are further amplified by the
The associate editor coordinating the review of this manuscript and relatively small number of public child speech corpora that
approving it for publication was Juan Wang . are available with useful annotations.

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
47628 VOLUME 10, 2022
R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

Current work done on TTS for child’s voices is limited. trend of data-hungry DNN-based TTS, TTS for children has
This is mainly due to the lack of child voice datasets and dif- practically been neglected due to the lack of large publicly
ficulty in creating such datasets. As TTS models require hun- available children’s speech datasets suitable for training such
dreds of hours of annotated data for training [2], performing networks. Prior to this DNN era, researchers worked on TTS
TTS for child voices can be quite challenging. The focus of for children using HMM-based models [3], [6].
this work is to explore the potential of state-of-the-art (SOTA) Collecting data for child speech research can be a chal-
TTS to build a pipeline for the synthesis of children’s voices lenging task. Most TTS datasets are created in studios with
with low data requirements. More specifically, if we can build expensive equipment: an adult will be using a microphone to
such a pipeline and demonstrate that it can reliably synthesize create a clean, noiseless, easy to understand, and meaningful
a useful number of distinct children’s voices, this pipeline audio. This task is not easy to produce and even more difficult
would enable the creation of large synthetic datasets that to implement with a child.
could further improve other aspects of child speech research One of the main differences between adult speech and child
such as automatic speech recognition (ASR), speaker recog- speech is the fundamental frequency. The pitch for children is
nition, etc. To better elaborate on this hypothesis, it is useful significantly higher than that of an adult [35]–[38]. The pitch
to review current SOTA in TTS technologies, followed by a for an adult voice lies between 70 to 250 Hz whereas the pitch
similar consideration for review in child speech research. for the children’s speech is between 200 to 500 Hz [39]. There
is also a difference in the speaking rate of children. It was
A. RELATED RESEARCH IN TTS noticed that average phoneme duration is longer in children,
Early research work on TTS synthesis can be traced therefore, leading to longer speaking rates as compared to
back to four/five decades ago when the task of TTS adult speech [38], [40]–[42]. The vocal tract of an adult is
was commonly tackled using concatenative and parametric larger as compared to children’s vocal tract and therefore
approaches [3]–[7]. Although these early methods were suc- produces different prosody features as compared to an adult
cessful in generating speech from text, they generally lacked voice [43], [44]. Hence, a substantial difference in children’s
naturalness. The audio generated using these approaches was voice characteristics and features can be seen as compared to
kind of muffled and sounded very robotic. an adult voice.
Recent state-of-the-art TTS models are largely based on Our work aims to solve the problem of TTS for children
deep neural networks (DNN) and can achieve more natural- using DNNs. To solve this problem, the huge challenge of
sounding/human-like synthesized speech. With the introduc- limited publicly available children’s speech datasets must
tion of Tacotron [8], a neural sequence to sequence the TTS first be overcome. To this end, this study considered the use
model, the quality of speech synthesis improved significantly. of an existing multi-speaker children’s speech dataset [45],
While there are newer approaches that are more efficient or which comes with an incomplete set of utterance transcrip-
use smaller models, etc., it is still representative of SOTA tions. In addition, this dataset has a lot of unusable data,
for the quality of the synthesized speech and is used as a such as empty/blank entries, extremely long entries as well as
benchmark for comparison with newer methods. Nonethe- inaccurate transcriptions. Firstly, the dataset is cleaned up to
less, Tacotron TTS is not very robust as it sometimes skips create a subset that is suitable for training a neural TTS model.
certain words and it also suffers from low inference speed Secondly, with the cleaned-up dataset, a multi-speaker TTS
[9]. Several methods have since been proposed to improve model is trained to generate synthetic speech for multiple
upon it such as Tacotron2 [10], FastSpeech [11], FastSpeech2 child speakers as a proof of concept for children’s TTS. The
[12], Transformer TTS [13], FlowTTS [14], GlowTTS [15], training involved fine-tuning an existing adult multi-speaker
etc. Similarly, there have been several improvements over TTS model [33] by way of transfer learning, with a few
the quality of synthesized waveforms by the introduction modifications as explained in later sections. This approach
of SOTA Vocoders such as WaveNet [16], WaveGlow [17], involves the training of a separate speaker verification model,
MelGAN [18], Hifi-Gan [19], WaveRNN [20], etc. These and it was preferred because it reduces the problem at hand
TTS models supported single speaker synthesis, but Deep- in two ways:
voice2 [21], introduced the use of speaker verification mod- 1) To train the speaker verification network, transcrip-
els [22]–[25] to achieve Multi-speaker TTS [26]–[34]. tions for the speech dataset are not required. Only the
speaker identities for the utterances are needed and it
B. CHILD SPEECH – LITERATURE AND CHALLENGES can also be trained on noisy speech without any nega-
While all SOTA TTS systems rely on large datasets to train, tive effects. This means that even the noisy children’s
the datasets mostly comprise speech taken from adult native speech dataset, which has incomplete transcriptions,
English speakers; hence, for low-resource languages and can be useful in training the verification model.
other target groups such as non-native adult speakers and 2) Being a transfer learning process, the pretrained TTS
child speakers, there remain challenges developing effective model can be finetuned sufficiently using the resulting
and suitable TTS models. Specifically, in comparison with cleaned set of children’s speech data.
adult TTS, child TTS has gained very little to no atten- Subjective and Objective Evaluation performed on the syn-
tion from the TTS research community. With the current thesized child voices confirms that the child voices generated

VOLUME 10, 2022 47629


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

TABLE 1. Dataset used in this work. TABLE 2. MyST dataset comparison [complete vs with transcript].

synthetically are very close to the real child voice in terms of


different acoustic features and MOS.
The rest of this paper is organized as follows. Section II
describes the methodology and datasets used in this study.
The experiments are presented in Section III, the result and
evaluation in Section IV, and finally, the conclusion and
future work in Section V.

II. PROPOSED METHODOLOGY


A. DATASETS USED IN THIS STUDY
The nature of this study, considering the challenge of limited
children’s speech datasets and the multi-step training process
involved, calls for the use of multiple large datasets, including transcripts are available is presented in Table 2. This table
adult speech datasets. All these datasets are described in provides information on the utterance count and duration of
Table 1. utterance concerning the duration range.
From Table 2, it was observed that 197.5 hours of child
• MyST [45]: My Science Tutor (MyST) children’s cor-
speech data is available with annotation. Although a lot of
pus consists of child speech collected using the inter-
this data can’t be used having different memory requirements
action of the student with a virtual science tutor. The
on different GPUs. In our experiments, that data between the
data consists of 393 hours of child speech collected from
range of 10-15 seconds to be most useful.
1371 students producing a total of 228,874 utterances.
Some initial experiments were performed on the MyST
45% of the data is transcribed at word-level leading to
dataset without using the Multi-speaker TTS approach (see
about 103,082 utterances, around 208 hours presented
section III.A). The results obtained from these experiments
in a .trn file format. The MyST corpus is used for this
were unintelligible. The output waveforms did not have any
paper because it is the biggest corpus of child speech
phonetic meaning and were missing quite some pronuncia-
freely available for research use.
tions. On a more detailed manual inspection of the MyST
• VoxCeleb1 [46] : VoxCeleb 1 contains audio recordings
dataset, a few common problems were identified. The tran-
of celebrity voices extracted from YouTube. It contains
scripts of some example audio files are listed below to illus-
153,516 utterances from 1,251 speakers.
trate the problems in the MyST dataset:
• LibriSpeech [47] : LibriSpeech is a read English speech
dataset derived from audiobooks. The data contains • Audio files containing noise in their utterances without
approximately 1000 hours of adult speech data from any phonetic meaning.
2400 speakers. The data is divided into two sets, ‘‘clean’’ • ‘‘<noise>’’
and ‘‘other’’ where the clean set contains less noisy data • ‘‘it’s glowing <breath>’’
as compared to the other set. The ‘‘clean’’ set contains • Audio files that are not coherent or indiscernible.
460 hours of data, and the ‘‘other’’ set contains 540 hours • ‘‘in oxygen right <indiscernible>’’
of data. • ‘‘can hear sound because of that <indiscernible>’’
• VCTK [48] : This dataset contains speech recordings
• Audio files are too small in length
from 110 English speakers each reading about 400 sen-
tences from a newspaper. The data contains recordings • ‘‘energy <noise>’’
from various English accents and is highly used in • Audio files are too long
multi-speaker TTS research. • ‘‘it’s trying to show us that all the things that it needs
all the things that the plants needs to grow it needs
1) PROBLEMS IDENTIFIED IN MYST DATASET soil on the bottom it needs at least a ground a top
A study on the MyST dataset was performed to measure the the a a top to lay on for the plant to grow so you
amount of data in MyST with and without a transcript. This can see it that’s only with flowers and plants it’s not
was done to extract data available with annotation and to see with vegetables and it needs and it needs the energy
if it can be used for training TTS. A comparison between from the sunlight to grow and it needs water because
the complete MyST dataset and filtered MyST dataset where somebody’s watering the plant.’’

47630 VOLUME 10, 2022


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

TABLE 3. MyST vs TinyMyST.

FIGURE 1. Model Overview: Speaker Encoder, Acoustic Model, and


Vocoder Models trained independently (from [33]).

not very accurate for the child speech and there were a
lot of mismatches between the transcripts and audio files.
This was probably due to fact that the pretrained forced
aligner was trained on adult speech and doesn’t work very
well for aligning child speech. Therefore, TinyMyST was
• Transcription containing text with no phonetic used (as described earlier) for performing all the child TTS
information. experiments.
• ‘‘(()) (()) (())’’
• Repetition of words/stammering noticed in children’s 3) DATA PREPROCESSING FOR TTS USAGE
voices. LibriSpeech and TinyMyST datasets were preprocessed as
• ‘‘um we measured how big a millimeter meter is a per the guidelines mentioned in LibriTTS [49]. The LibriTTS
meter and a kolome- a ∗ kilometer ∗’’ dataset was specifically created for TTS research, therefore
Our examination of MyST led us to further clean the MyST similar guidelines were followed in our experiments. The
dataset for TTS training. In this process a subset of MyST, following changes were made:
hereafter referred to as TinyMyST was created. • Audio files were converted to 16-bit depth audio files
with 24Khz sampling rate (WAV format), This was done
2) TINYMYST using the pydub2 audio library.
It is a small subset of the MyST dataset created using var- • Text data was normalized by replacing abbreviations and
ious pre-processing scripts to make the data suitable for punctuations.
TTS acoustic model training. MyST was cleaned to select • Whitespaces were normalized
only audio files with existing transcriptions. All audio files • All characters were made uppercase.
lesser than 10 seconds and greater than 15 seconds were
removed. The utterances shorter than 10 seconds contained B. MULTI-SPEAKER TTS MODEL
mostly noise or unintelligible speech and those longer than
The neural network used to achieve TTS for children in this
15 seconds were removed to avoid GPU memory overflow
study is based on [33], It works by combining a speaker verifi-
during training. All the transcript files were converted from
cation network with the SOTA Tacotron TTS model. Though
.trn format to .txt file format.
Tacotron is SOTA for TTS, it was designed to be trained using
The TinyMyST dataset still contains a lot of noisy data.
a single-speaker speech dataset such as the LJSpeech [50]
Some of the excessively noisy data were removed man-
dataset, hence, it can only synthesize speech with acoustic
ually by inspecting the transcripts and listening to the
characteristics of the single speaker whose data was used
audio samples. The data obtained after cleaning contained
in training. To function effectively for multiple speakers,
7152 utterances and accounted for 19.22 hours. A detailed
Tacotron needs to be adapted for that purpose. This adaptation
comparison of MyST and TinyMyST datasets was performed
has been achieved in this multi-speaker TTS model [33] by
to see differences in the two datasets in terms of speaker iden-
introducing different speaker identities in the form of speaker
tities and utterances (see Table 3). TinyMyST dataset on aver-
embeddings as additional input to the Tacotron network. As a
age contained 1.72 minutes per speaker having 670 speakers.
result, the multi-speaker TTS [33] comprises three different
Speaker identity ‘013023’ had the most data with 8.77 min-
neural network models, each of which focuses on a specific
utes and speaker identity ‘018216’ has the least data with
subtask namely, Speaker Encoder used for speaker verifica-
10.01 seconds. The speaker ‘013023’ had the most data in
tion task, Acoustic model used for spectrogram synthesis,
MyST as well to be around 110 minutes.
and a Vocoder for audio waveform generation (as shown in
To extract more TTS usable data, an audio sample from
Figure 1).
more than 15 seconds long can be used to split them into
For our work, generalized end-to-end (GE2E) loss was
smaller chunks. A forced aligner1 is used to align the audio
used for speaker verification [22], Tacotron1 as an acoustic
files with transcripts. Time alignment information from the
model [8], and WaveRNN as Vocoder [20]. The original
alignments to split the longer audio files into smaller sam-
approach [33] is adapted for child speech synthesis by first
ples, however, it was observed that the audio alignment was
1 [Link] 2 [Link]

VOLUME 10, 2022 47631


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

pretraining the model on an adult speech dataset after which, TABLE 4. Speaker encoder training details.
it is fine-tuned with the child speech dataset.
The speaker encoder generates speaker embeddings,
encoding speaker identity information extracted from the
utterances. Similar voices are mapped closer to each other
in a latent space representation. The acoustic model gen-
erates spectrograms from text conditioned on the speaker
embeddings. The vocoder then converts these spectrograms
into audio waveforms. At inference time, a short refer-
ence utterance (ground truth) of a child’s voice is passed
through the speaker encoder to generate the corresponding
speaker embeddings, on which the acoustic model will be
conditioned. The three different neural network models are
described as follows.

1) SPEAKER ENCODER
The first stage of the multi-speaker TTS training involves the
training of a speaker verification (speaker encoder) model.
Speaker Verification is the process of determining if an utter-
ance belongs to a specific speaker. The speaker encoder is
used to train the model for the speaker verification task using FIGURE 2. Pipeline for Speaker Encoder training. The dotted line
represents the training loop for the Speaker Encoder training.
a mix of noisy and clean speech data without transcripts.
The data used consists of both adult and child speech data
from thousands of speakers (see Table 4). This was done to relatively insignificant improvements were seen in the EER
introduce both child and adult speakers in the model for better after this point.
generalization. The output of this model conditions the acous- All the datasets were pre-processed into the coding format
tic model to generate the required mel-spectrograms from required for training the encoder as described in [51]. Even
a reference speech signal of the target speaker. The model though half the MyST dataset is not transcribed, the complete
is trained to capture the characteristic features of different MyST dataset can be used for Speaker Encoder training as it
speakers. does not require any transcription data. The pipeline for the
The model takes input as log mel-spectrograms computed speaker encoder training can be seen in Figure 2.
from utterances of each speaker, trains using the GE2E loss A UMAP projection [52] is created to visualize the training
and converts them into a fixed dimensional vector called by taking a random set of 10 utterances from 10 speakers.
d-vectors. These d-vectors are optimized over GE2E loss Utterances with similar embeddings are located close to each
to differentiate the speakers, such that the same speakers other in the latent space representation and have similar
have embeddings with high cosine similarity and different speaker characteristics.
speakers are far apart in the embedding space. This model creates individual clusters of speaker embed-
During training, complete utterances are segmented into dings as can be seen in the UMAP projection (see Figure 3.
partial utterances of 1.6 seconds. These parameters were kept Each point on UMAP represents an utterance. The same color
the same as explained by authors [51], [22]. The utterance points represent the same speaker. Encoder gradually learns
embedding is calculated using 800ms windows for inference, to separate the speakers. Initially, there is a lot of overlap
with a 50% overlap. The silence was removed from the utter- across speakers, but eventually, each speaker has their utter-
ances using the webrtcvad3 tool for Voice Activity Detection ances clustered and well separated from the other speakers.
(VAD). Each segment is passed through the network individ- The training evolves with increased training steps.
ually, the outputs are averaged and normalized to create the
final utterance embedding as described in [22]. 2) TACOTRON ACOUSTIC MODEL
The encoder model is trained using 4 datasets, MyST, Vox- For the speech spectrogram synthesis, the TTS model archi-
Celeb1, LibriSpeech, and VCTK. Equal Error Rate (EER) is tecture and hyperparameters used in this study are the
used as a metric for the validation of the speaker encoder. The same as in the work of [51] (More details are provided
default EER metric from [51] is used in this work as authors in Section III). The authors used a modified version of the
of [33] have not explicitly specified the training, test, and original Tacotron architecture [8]. The model consists of an
validation criterion they are using for EER calculation. The encoder, an attention-based decoder, and a post-processing
EER values are presented in Table 4. The model trained for network. Since Tacotron is originally a single-speaker TTS
one million steps was used in the multi-speaker TTS model as model, it was modified to work for multi-speaker TTS by con-
necting the speaker encoder to it. Speaker embeddings from
3 [Link] the encoder are concatenated with text (character/phoneme)

47632 VOLUME 10, 2022


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

FIGURE 3. UMAP projections at different training steps for speaker


encoder training. Ten different colors represent ten different speakers FIGURE 4. Pipeline for Acoustic Model training. A model with solid
with ten utterances each. contour represents the pretrained model. Dotted contours represent the
acoustic model training loop. Fine-tuning step for Acoustic model
training. Acoustic model training I represent the acoustic model being
embeddings from the text encoder, after which an attention trained with LibriSpeech dataset for up to 250k iterations. Acoustic model
mechanism is applied prior to decoding into a spectrogram. training II represents fine-tuning the acoustic model I with the TinyMyST
dataset from 250k iteration onwards up to 750k iteration.
Unlike the speaker encoder, the acoustic model takes in both
audio(utterance) and associated text(transcript) as inputs.
In this work, the acoustic model was first trained with only transformations in four-way connections. These connections
adult speech data (acoustic model training I), specifically, are concatenated at different steps to generate the correspond-
the Librispeech ‘clean’ data, until it started to converge at ing vector representation. This vector is passed through two
250k steps and then finetuned with the TinyMyST child dense layer connections which finally generate the encoding
speech dataset (acoustic model training II) for up to 750k of raw audio. The output audio is generated at a 16-bit depth
additional steps (more details in Section III). The pipeline for and 16 khz sampling rate.
the acoustic model training can be seen in Figure 4. The predicted mels from the acoustic model trained on
LibriSpeech (from acoustic model training I) were used to
3) WAVERNN VOCODER train the vocoder. The vocoder trained up to 250k iterations
The vocoder used is WaveRNN [20], which is an improve- was used to generate all waveforms in this study. The pipeline
ment over the WaveNet [16] originally used by the authors for vocoder training can be seen in Figure 5. Fine-tuning
of [33]. WaveRNN is a recurrent network for perform- experiments with the TinyMyST dataset didn’t improve the
ing sequential modeling of audio from mel-spectrograms. quality of the vocoder (more discussion in Future work).
An alternative version of WaveRNN is used, having a few Vocoder for child TTS hasn’t been explored before in
architectural changes as provided by the author in [53] due detail. This is a new area of research. It was observed
to the popularity of the model as it reduces sampling time that WaveRNN has popularly been used as a universal
while maintaining high output quality. WaveRNN uses Gated vocoder [54]–[56] and it evidently works well with unseen
Recurrent Unit (GRU) in comparison to convolutions used speakers in multi-speaker models as well [57]. Therefore, for
in WaveNet. The input mel-spectrograms and their corre- the scope of this paper, WaveRNN (trained on LibriSpeech)
sponding waveforms are segmented at each timestamp. A 1D is used as a universal vocoder with synthetic child voices.
Resnet-like model is used to generate features for layered
connections in the alternative WaveRNN architecture. The III. EXPERIMENTS
upsampling is also performed on the mel-spectrogram to A. INITIAL EXPERIMENTS
match the length of the target waveform. The resulting vector In our initial experiments, multiple SOTA TTS models
is passed through a combination of GRU and dense layer [10], [13]–[15], [21] were unsuccessfully trained, including

VOLUME 10, 2022 47633


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

FIGURE 6. Alignment plot for Tacotron 2 trained with MyST dataset for up
to 200k steps.

FIGURE 5. Pipeline for Vocoder training. Models with solid contours are
pretrained models. The dotted contour represents the training loop for
Vocoder.

Tacotron 2,4 using the transcribed subset of the MyST dataset.


Figure 6 shows an example of an alignment plot from
Tacotron 2 training. As can be seen, there was no sign of
alignment even after 200k iterations.
FIGURE 7. Alignment plot for Tacotron 2 trained up to 200k steps with
Further experiments were conducted using the cleaned sub- TinyMyST Dataset.
set of MyST (TinyMyST), which showed some alignments as
seen in the Tacotron 2 alignment plot in Figure 7. However,
though child-like in terms of pitch, the synthesized speech
signals were completely unintelligible. Missing information
such as ‘End of sentence’ was observed which mostly con-
tained noise content.
Next, fine-tuning the pretrained NVIDIA Tacotron 2 model
on a single child’s MyST dataset utterances resulted in
slightly intelligible but highly robotic and unnatural synthe-
sized speech. Figure 8 shows the improved alignment plot
from the finetuned Tacotron2.
Since there were not enough MyST utterances for a single FIGURE 8. Alignment plot for Tacotron 2 trained up to 200k steps with
child to sufficiently train Tacotron2, different multi-speaker TinyMyST Dataset, pretrained with LJ Speech Dataset up to 100k steps.
TTS models were explored [21], [26], [27], [58] and the
speaker verification-based method [33] produced the most Additional parameters settings are mentioned here.5 The
promising results. Hence, this method was used in our main default embedding size of 256 was used for this training.
experiments. For the acoustic model, the network was trained using
a learning rate of 0.0001 for 250K steps (pretraining) and
0.00001 for 750k steps (fine-tuning). The batch size was kept
B. MAIN EXPERIMENTS
constant at 72. Entire training (up to 750k steps) took 9 days to
As seen in our methodology (Section II.B), a modified
complete. Additional parameters details were kept the same
approach based on [33] was used by incorporating an extra
as Tacotron 1, these details are mentioned here.6
layer of fine-tuning in the training step.
The alignments plot for encoder-decoder timestamps can
The proposed neural child voice TTS was trained on a Tesla
be seen in Figure 9, the x-axis represents the encoder
V100 GPU. Each of the three networks – Speaker Encoder,
timesteps and the y-axis represents the decoder timesteps
Acoustic model, and Vocoder were trained separately.
of Tacotron training. The training is done on LibriSpeech
The Speaker Encoder was trained with a batch size
up to 250k steps generated a good alignment plot. Align-
of 128 and a learning rate of 0.0001. The model was
ment weakens when switched to TinyMyST Dataset, but it
trained for 15 days for up to 1M steps. EER of 5% was
observed at this point with no further improvement afterward. 5 Encoder Hyperparameters: [Link]
Voice-Cloning/blob/master/encoder/params_model.py
6 Acoustic Model Hyperparameters: [Link]
4 NVIDIA/tacotron2: [Link] Time-Voice-Cloning/blob/master/synthesizer/[Link]

47634 VOLUME 10, 2022


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

There are many objective and subjective evaluation meth-


ods proposed by researchers [60]–[66]. These traditional
speech evaluation methods work well for evaluating adult
speech but are not so suitable for child speech. A perfect adult
speech will contain fluent pronunciation of a word/phoneme
however this is not the case for most child speech. Natural-
ness in child speech includes pauses, breaks, and pronunci-
ation difficulties in the speech. Other challenges were noted
with the start and end of phrases where children tend to be
somewhat hesitant when starting a phrase and may wander
towards the end of one. Children can also mispronounce
words, or struggle with the phonetics of a particular phrase.
These characteristics tend to manifest in the speech model
and a range of artifacts were noted that affect the quality of
the phrases synthesized by our pipeline. It was noted that
the first or last words in many phrases were either missed
entirely or subject to various distortions or artifacts. In the
FIGURE 9. Alignment plots at different training steps during transfer middle of a phrase, there could occasionally be slurring or
learning from adult to child Tacotron TTS. arbitrary elongating of one or more words. Another artifact
observed was that the pace or tone of voice could change
gradually improves with increasing training steps. During abruptly in the middle of a phrase. Despite these artifacts,
inference, our model was tested on multiple checkpoints the majority of phrases were quite intelligible, and a large
taken at intervals of 50k iterations. A few of these iterations proportion was also very natural sounding. Therefore, there
are mentioned in Figure 9. Even though the alignment at is a need for a better subjective evaluation method for child
some of these steps looks the same, an improvement was speech synthesis.
noted over time with the synthesized child voices. This was In the following, we present the results obtained using
determined subjectively during training by listening to the the proposed subjective evaluation method and the various
synthetic child speech generated. The training was halted listening tests performed (subsection 4.A), two objective
at 750k steps as improvements in the alignment graph had evaluation methods, based on MOSNet (subsection 4.B) and
become imperceptible after 700k steps. The output waveform an ASR system (subsection 4.D) and we evaluate the simi-
did not show any improvements beyond this step. The model larity of the synthesized speech and natural child speech (in
trained up to 750k iterations is used to provide audio samples subsection 4.C).
in this paper.
The Vocoder was trained at a batch size of 128 and learning A. PROPOSED SUBJECTIVE EVALUATION METHOD
rate of 0.0001 and took 4 days of training to reach 250k To check the phonetic coverage of our child speech TTS,
iterations. Most of the parameters for the vocoder were kept Harvard sentences [60] were used, which are a set of 720
the same as the original code.7 phonetically balanced sentences. These sentences cover most
The synthetic child voices during inference were natural of the phoneme range and were designed to be implemented
sounding and the trained model demonstrated an ability to with Voice over Internet Protocol (VoIP) technology. These
synthesize quite challenging phrases that were unseen in the texts were used to generate synthetic child speech. This was
TinyMyST dataset. This was tested by using ‘tongue twisters’ done to check the subjective quality of synthesized audio with
as a reference text for synthesizing speech. However, it was respect to phoneme coverage.
also noted that some phonemes were not synthesized cor- Our evaluation method uses a MOS-like evaluation with
rectly and lost their meaning during synthesis. These findings different categories for scoring. When generating synthetic
are discussed in more detail in Section IV. voices using Harvard sentences, it was observed that some
Code-related material and synthesized speech from sets of phonemes were not pronounced correctly even when
these experiments will be made available in our GitHub synthesized using different reference child speakers (more
Repository.8 detail in a later section). After our initial subjective study of
these 720 synthesized audio samples, it was decided that a
IV. RESULTS AND EVALUATION more detailed evaluation protocol was required to address the
The evaluation in TTS is usually done by taking a Mean various artifacts observed and identify what additional data
Opinion Score (MOS) [59] on the synthetic speech for Speech samples might be needed to further improve our model. For
Similarity and Speech Naturalness. this reason, our evaluation was performed in two phases. For
7 Vocoder Hyperparameters: [Link] each of the two phases, different evaluators were gathered to
Voice-Cloning/blob/master/vocoder/[Link] perform the speech evaluation. Each evaluator was asked to
8 GitHub for this paper: [Link] listen to synthetic audio files using Headphones/Earphones

VOLUME 10, 2022 47635


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

in a noise-free environment. They were asked to rate each TABLE 5. MOS (from 1 to 5) explained for speech intelligibility and voice
naturalness.
of the synthetic voices assigned to them from a range of
1 to 5, for each of the different categories in two phases. The
categories included Speech Intelligibility, Voice Naturalness,
and Voice Consistency. Voice Consistency contained three
sub-categories of its own namely, Start of Phrase quality,
Middle of Phrase quality, and End of Phrase Quality.
Evaluation data was provided in a OneDrive Environment.
All the synthetic voices were shared in a common OneDrive
folder to the evaluators and a common spreadsheet was circu-
lated containing the utterance ID of Harvard sentences used
for synthesizing a child’s voice. While listening to many
different natural child voices, it was also noticed that recorded
child audio can be a difficult task to understand if not pro-
vided with a suitable transcript. Some of the child’s speech
can be non-meaningful as mentioned in problems with the
MyST section. After performing many different tests and
trials using child speech, the use of transcripts as a part of TABLE 6. MOS from phase-I evaluation with 95% confidence interval
MOS-based evaluation is considered to be a more natural
way of evaluating child speech. Therefore, corresponding
transcript information is also provided in the spreadsheet to
each evaluator to base their conclusion on ‘what they hear
in child audio’ and ‘what they read in child transcripts’.
This way more coherent patterns can be observed among the The spreadsheet was later analyzed to get the final mean
phonemes and graphemes in a child’s voice for each of the opinion score in each category. MOS of 3.88 for voice nat-
mentioned categories. An example of this spreadsheet can be uralness and 4.13 for speech intelligibility was observed as
seen here.9 seen in Table 6.
By performing the evaluation using OneDrive environ- An average score for each of the 720 sentences was cal-
ment, it was easy to distribute the synthetic speech files to dif- culated for the combined value of speech intelligibility and
ferent evaluators without having to spend time and resources voice naturalness. All the 720 sentences were sorted into
on expensive Mushra-based evaluations [67] or crowdsourc- difficult and easy sentences with respect to the children’s
ing the evaluation task on platforms like Amazon Mechani- linguistic capabilities. This was done to keep track of Harvard
cal Turk (AMT) [59], [68]. Mushra-based evaluations were sentences where synthesized speech becomes unintelligible
also avoided due to potential biases that can occur in these and inarticulate.
tests and how these biases can impact synthetic child voice
evaluation for MOS [69]. Most of these TTS evaluations 2) PHASE-II EVALUATION
have been conducted before with synthetic adult speech, this After our phase-I evaluation, a common set of sentences
novel synthetic evaluation is implemented for first-time with were observed where pronunciation sounds unintelligible
synthetic child speech. Using a common spreadsheet made it at the start, middle, or end of sentences for specific
effective to perform analysis of spreadsheet for MOS using words/phonemes. There was an inconsistency in voice qual-
pandas and other python-based tools. ity. These sets of sentences are the ones that were not learned
properly during training or were missing in the training
1) PHASE-I EVALUATION dataset for child audio. To make a note of these sentences,
For the phase-I evaluation, all 720 Harvard sentences were extra categories of ‘Voice Consistency’ were added to the
generated using our proposed TTS method. Two random ref- phase-I evaluation. Therefore, all the 3 sub-categories under
erence utterances were selected from the TinyMyST dataset Voice Consistency were used in the second phase of the eval-
and were used to generate all the Harvard sentences. These uation. These subcategories included ‘Start of Phrase Qual-
720 sentences were shared among 5 evaluators in a spread- ity’, ‘Middle of phrase quality’ and ‘End of phrase quality’.
sheet document, who rated the voices from 1 to 5 based The MOS ratings from 1 to 5 for each of these categories were
on Speech Intelligibility and Voice Naturalness. The MOS also explained in the spreadsheet as mentioned in Table 7.
ratings from 1 to 5 were further explained in the spreadsheet For the second phase of evaluation, the evaluation was
file as can be seen in Table 5. undertaken by 20 evaluators divided into 4 groups. This was
done as per the guidelines mentioned in [70] for performing
MOS evaluations. For each group, a speaker identity was
9 Example Spreadsheet: [Link] selected from the TinyMyST dataset. All the speaker iden-
blob/main/synthetic%20evaluation%20example/[Link] tities were sorted, and the top 20 speaker identities were

47636 VOLUME 10, 2022


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

TABLE 7. MOS (from 1 to 5) explained for voice consistency and its three TABLE 9. MOS from phase-II evaluation with 95% confidence interval.
sub-categories.

room for improvement in the ‘End of phrase quality’ of Har-


vard sentences. There is information loss observed at the end
of most sentences containing inarticulate and unintelligent
information or noise. The reason for this information loss can
be redirected back to the child dataset used for training. Even
though TinyMyST is much cleaner than the MyST dataset,
it still contains some of the problems seen in Section II.A.1.
The information obtained from voice consistency will be
discussed more in future work.
A similar experiment was also performed using the real
utterances from Table 8 to obtain a baseline MOS on natural
child speech. 15 random real utterances were selected from
TABLE 8. Selected speaker identity information in TinyMyST VS TTS
utterances for the same speakers. the real speakers mentioned in Table 8. Evaluators were asked
to perform a similar evaluation as done in phase-II evaluation
for all the selected 60 utterances. A comparison between the
baseline MOS on Natural MyST and synthetically generated
utterances is mentioned in Table 10.
Synthetic Speech MOS for three categories is very close to
Natural Speech MOS. There is a MOS difference of ‘0.26’
for Speech Intelligibility, ‘0.16’ for Voice Naturalness, and
selected, having the most minutes. Among these 20 identities, ‘0.12’ for Voice Consistency between natural and synthetic
4 speaker identities were randomly selected. All the 4 groups speech. From Table 10, it can be concluded that the MOS
are named as ‘013020’, ‘008045’, ‘002113’, and ‘995737’, for Natural and Synthetic child speech are quite close to
corresponding to each identity label. This approach was taken each other. This subjective evaluation approach is proposed
to select speakers with the most data and also to keep the as a part of this paper. Due to very limited work done on
process randomized. More information on these selected child speech synthesis, we did not find any reliable way of
speakers can be seen in Table 8. This table is also used for performing subjective evaluation over the synthesized child
speaker similarity and objective intelligibility experiments in speech. From our experience with the evaluation of synthetic
the future sections. child speech, this new metric of evaluation can help evalu-
A reference child utterance was selected randomly from ate synthetic child speech and can help further this area of
each of these groups, and 50 Harvard sentences were selected research. It is also intended to use this proposed approach for
randomly for each of the groups. Therefore, 50 Synthetic our future work with child speech synthesis.
utterances were generated, and all the evaluators were asked
to rate the utterances assigned to them. B. OBJECTIVE NATURALNESS EVALUATION USING A
MOS results from the phase-II evaluation are presented PRETRAINED MOSNET
in Table 9. MOS of 3.95 was observed for Speech Intel- For this objective evaluation, a pretrained MOSNet was used,
ligibility, 3.89 for Voice Naturalness, and 3.96 for over- which is trained on VCC 2018 dataset from Blizzard Chal-
all Voice Consistency (including the three sub-categories). lenge [66] comprising of adult speech. According to their
MOS of 4.07 was observed for ‘Start of phrase quality’, paper, MOSNet predictions yield a high correlation to human
4.18 for ‘Middle of phrase quality’, and 3.62 for ‘End of ratings. As MOSNet was trained on adult speech, it is unlikely
phrase quality’. The MOS score implies that the quality of that it will generalize well for child speech. It won’t be
synthesized child speech is quite good. However, there is still possible to train a MOSNet with child voices as there is not

VOLUME 10, 2022 47637


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

TABLE 10. MOS natural speech VS MOS synthetic speech with 95% TABLE 11. MOSNet output for 5 samples.
confidence interval.

TABLE 12. MOSNet output for 50 samples with 95% confidence interval.

reference child audio and synthetic child audio. This gave us


a comparison between reference and synthetic child voices
as to how close they are to each other in terms of audio
features calculated using MOSNet. The results confirmed that
MOSNet output for reference child speech and synthetic child
speech are very close to each other with a comparative MOS
difference of 0.3.

C. SPEAKER SIMILARITY EVALUATION USING A SPEAKER


VERIFICATION SYSTEM
Speaker similarity between a synthesized speech and a real
speech can be calculated using a Speaker Verification (SV)
system. The pretrained speaker encoder from section 2.B.1.
was used with a third-party tool10 to extract and visualize
the speaker embeddings. This tool uses cosine distance to
calculate the similarity between the two embeddings. The
same speakers mentioned in our subjective evaluation (see
Table 8) were used for this evaluation. 10 utterances were
randomly selected for both real and synthetic speech for each
of the 4 speakers mentioned in Table 8. 1 male and 1 female
speaker from the LibriSpeech dataset were also added with
FIGURE 10. Spectrogram comparison between reference and synthesized 10 utterances each to show the speaker similarity comparison
child audio for 5 audio samples used with MOSNet.
between an adult and child speaker. A visualization of this
similarity in a 2D projection can be seen in Figure 11, ‘gt’ is
enough data to perform a large-scale evaluation such as a used as a label for the ground truth of the speaker and ‘ss’ is
blizzard challenge. This objective evaluation was performed used as a label for the synthetic speech of the same speaker.
to see the correlation between reference child audio and ‘Adult_Male’ and ‘Adult_Female’ are two randomly selected
synthetic child audio. A random set of 50 utterances were male and female speakers from the LibriSpeech Dataset.
selected from the TinyMyST dataset as a part of this inside From Figure 11, it can be inferred that Male, Female, and
test. These utterances were used as reference utterances and Child speech have a difference in similarity from each other.
the corresponding transcripts were used to generate synthetic Male and Female adult speakers are far apart from each other
speech for each of these utterances. This gave us 50 reference and from child speakers in this 2D projection of speaker
and 50 synthetic utterances which were used to calculate embeddings.
MOS using MOSNet. MOS score for 5 samples can be seen To further comment on the similarity between real child
in Table 11. The spectrograms for these 5 samples can be seen speech and synthetic child speech, the ‘child speech’ contour
in Figure 10. from Figure 11 is extended to get a more visual representation
Table 12 shows the overall MOS output for MOSNet. MOS of embeddings. This can be seen in Figure 12. The ‘gt’ labels
of 2.96 was observed for reference child audio and 2.66 and very close to ‘ss’ labels in this 2D projection space.
for synthetic child audio. There is only a 0.3 difference in These embeddings are 256-dimensional feature vectors
MOS between reference and synthetic child voices. MOSNet trained by our speaker encoder. Therefore, cosine similarity
trained on adult speech data is not expected to give MOS was used to further calculate the cross-similarity between
ratings correlated with human MOS ratings for child speech
data. MOSNet was only used to get a correlation between the 10 Resemblyzer: [Link]

47638 VOLUME 10, 2022


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

FIGURE 11. Projections of embeddings between different real and


synthetic child speech along with adult speech. The child Speech region
[both ground truth and synthetic speech] is outlined by a solid black
rectangle. The projections include a cluster of 10 voices selected from
10 different speakers. ‘ss’ refers to synthetic child speech and ‘gt’ refers
to ground truth child speech.

FIGURE 13. Cross-similarity between 10 speakers in Set A and Set B. The


rectangular black box represents the similarity between real and
synthetic child speech for respective speakers in set-A and set-B. Set-A is
along the x-axis and Set-B is along the y-axis. ‘ss’ represents the synthetic
speech and ‘gt’ represents the ground truth (real) speech.

speaker similarities between real child speech and synthetic


child speech and to draw a conclusion that our synthetically
generated child speech is very close to real speech in terms
of speaker similarity with an average similarity of 81%.
FIGURE 12. Projections of embeddings between different real and
synthetic child speech. A solid black line is used to show the distance
between the ground truth and synthetic speech from the same speakers. D. OBJECTIVE INTELLIGIBILITY EVALUATION USING A
This line was drawn from the centroid of each cluster to show the visual PRETRAINED ASR SYSTEM
representation of similarity between real and synthetic speech.
A pretrained wav2vec2 model is used to provide verification
on synthetic utterances. A comparison of the speech tran-
each speaker. Each of the 10 Speakers with 10 utterances
scription between real and synthetic child voices is presented.
each (1 Adult Male, 1 Adult Female, 4 Ground Truth Child,
Child speech recognition is a challenging task of its own. The
and 4 Synthetic Speech Child) were divided into 2 sets A and
ASR on child speech is a part of our future work. Our intent
B. Embeddings are extracted for each of the utterances for
to use this model for this paper is based on the popularity of
each of the sets and averaged together for each speaker. This
the model, being SOTA on adult speech. A wav2vec2 model
gave us 10 unique speaker embeddings in sets A and B each
trained on adult speech data is used to provide that compari-
for 10 speakers. Cosine similarity is finally used to measure
son. This speech transcription was obtained for the synthetic
the similarity between each of the 10 speaker embeddings in
and real utterances mentioned in Table 8. A random set of 30
sets A and B. A plot for the cross similarity between speakers
utterances for each speaker for both real and synthetic voices
can be seen in Figure 12.
are selected. The instruction for using this model is mentioned
In Figure 13, speaker similarity between synthetic speech
in their Github.11
and ground truth for speaker ‘995737’ is 0.91, whereas for
A comparison of this model is also provided using adult
speakers ‘013020’ and ‘008045’ is approximately around
speech by selecting the equal number of adult voices from the
0.82 and finally for the speaker ‘002113’ is approximately
LibriSpeech dataset. Word Error Rate (WER) is calculated
0.7. This cross-similarity matrix gives us an idea of how
from the output of wav2vec2 and is mentioned in Table 13.
close synthetic child voices are in comparison to the real
The Flashlight12 library is used to calculate the WER using
child voice. It also shows us how different an Adult Male
Viterbi decoding. No external language model (LM) was
and Female Speech is in comparison to a child’s speech.
used.
Overall, the similarity between most of the child and adult
From Table 13, it can be inferred that the WER for Adult
speech is between 0.3-0.4 whereas the similarity between
Speech (Librispeech_test_clean) is 3.43, evidently, due to the
most of the synthetic child speech and ground truth child
model being trained on adult speech data, the WER for real
speech is between a range of 0.65-0.85. Cross-similarity
child speech is 15.27 and in comparison, WER for synthetic
across the diagonal signifies that an utterance in set-A is 95%
utterances is 25.63. An ASR model was able to recognize
similar to utterances in set-B having the same speakers. More
conclusions can be drawn from Figure 13, however, for the 11 wav2vec2: [Link]
scope of this research, it is only used to show the different 12 Flashlight: [Link]

VOLUME 10, 2022 47639


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

TABLE 13. WER on adult speech, real child speech and synthetic child to increase the training dataset. The information collected
speech.
from our subjective evaluation such as voice consistency in
Harvard sentences will be used to improve child speech.
This information will be used to collect better TTS-based
child speech data based on Harvard sentences to accord with
‘end of the phrase’ information loss and voice inconsistency
75% of the synthetic speech with a relative difference of observed with our current results. The use of synthetically
10 WER when compared with real child speech recognized generated child speech to improve other areas of child speech
by the same model for the same speakers. research such as ASR and speaker recognition will also be
investigated in future work. TTS-generated child voices can
V. CONCLUSION AND FUTURE WORK be used as a data augmentation technique for training these
In this paper, a pipeline for generating synthetic child speech models with additional data. It is also intended to use the
in a limited training data scenario is proposed. A small subjective evaluation method proposed in this paper for per-
set of child speech data is created by cleaning an existing forming all future subjective evaluations with TTS generated
child speech dataset and making it suitable for TTS training. child speech.
A transfer learning approach is used to train our model with
adult speech data in a pretraining setting and child speech data ACKNOWLEDGMENT
as low as 19 hours for fine-tuning. MOSNet based objective The authors would like to thank experts from Xperi-Ireland:
evaluation shows a high correlation between real and synthe- Gabriel Costache, George Sterpu, and the rest of the
sized child voices. A subjective evaluation method suitable team members for providing their expertise and feedback
for synthesized child speech is also proposed and demon- throughout.
strated. Subjective MOS of synthesized voices is observed
as 3.95 for speech intelligibility, 3.89 for voice naturalness, REFERENCES
and 3.96 for overall voice consistency which is very close [1] J. Xu, X. Tan, Y. Ren, T. Qin, J. Li, S. Zhao, and T. Y. Liu, ‘‘LRSpeech:
Extremely low-resource speech synthesis and recognition,’’ in Proc. 26th
to Natural speech MOS. These MOS values tell us about ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, Assoc. Comput.
how good the synthesized child voices are. However, voice Mach. New York, NY, USA, 2020, pp. 2802–2812, doi: 10.1145/3394486.
inconsistency for ‘End of phrase quality’ containing noise and [2] K. R. Prajwal and C V Jawahar, ‘‘Data-efficient training strategies for
neural TTS systems,’’ in Proc. 8th ACM IKDD CODS, 26th COMAD.
unintelligible information was also observed. There is scope New York, NY, USA: Association for Computing Machinery, 2021,
for improvement for these phrases. WER for synthetic child pp. 223–227, doi: 10.1145/3430984.3431034.
voices using a pretrained adult speech wav2vec2 ASR model [3] O. Watts, J. Yamagishi, K. Berkling, and S. King, ‘‘HMM-based synthesis
of child speech,’’ in Proc. 1st Work. Child, Comput. Interact. (ICMI Post-
came to be 25.63 as compared to WER of real child voices Conf. Work., 2008.
of 15.27. Synthetic child speech samples can be viewed in [4] H. Zen, K. Tokuda, and A. W. Black, ‘‘Statistical parametric speech
our GitHub repository.13 Multi-speaker TTS can be the key synthesis,’’ Speech Commun., vol. 51, no. 11, pp. 1039–1064, Nov. 2009,
doi: 10.1016/[Link].2009.04.004.
to child speech synthesis with limited training data. Child [5] P. K. Muthukumar and A. W. Black, ‘‘A deep learning approach to
speakers with speech duration between 5-7 minutes in TTS data-driven parameterizations for statistical parametric speech synthesis,’’
training gave 81% average cosine similarity with a synthetic Carnegie Mellon University Pittsburgh, Pittsburgh, NA, USA, 2014.
[6] O. Watts, J. Yamagishi, S. King, and K. Berkling, ‘‘Synthesis of child
speech from the same speakers. This choice of the model speech with HMM adaptation and voice conversion,’’ IEEE Trans. Audio,
allows the TTS to learn useful speaker information which can Speech Language Process., vol. 18, no. 5, pp. 1005–1016, Jul. 2010, doi:
be leveraged to produce better quality synthetic voices even 10.1109/TASL.2009.2035029.
[7] R. Maia, H. Zen, and M. J. F. Gales, ‘‘Statistical parametric speech synthe-
with limited child speech. sis with joint estimation of acoustic and excitation model parameters,’’ in
For future work, our aim is to improve this method by Proc. SSW, 2010, pp. 88–93.
incorporating more information to our multi-speaker TTS [8] Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly,
Z. Yang, Y. Xiao, Z. Chen, S. Bengio, and Q. Le, ‘‘Tacotron: Towards end-
model such as duration predictor and energy as implemented to-end speech synthesis,’’ 2017, arxiv:1703.10135.
in FastSpeech2 [12]. The trained vocoder was also finetuned [9] J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu,
on the TinyMyST dataset. However, there was no significant ‘‘Non-attentive tacotron: Robust and controllable neural TTS synthesis
including unsupervised duration modeling,’’ 2020, arxiv:2010.04301.
improvement in the quality of the generated audio waveforms [10] J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen,
and an additional noise was observed in some of the synthesis. Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis,
More child speech data would be required to achieve any and Y. Wu, ‘‘Natural TTS synthesis by conditioning wavenet on MEL
spectrogram predictions,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal
significant improvement over the quality of the vocoder. It is Process. (ICASSP), Apr. 2018, pp. 2756–2761.
also intended to implement GAN-based SOTA Vocoders such [11] Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Y. Liu,
as HiFi-GAN [19] for future experiments. More experiments ‘‘FastSpeech: Fast, robust and controllable text to speech,’’ in Proc. Adv.
Neural Inf. Process. Syst., 2019, vol. 32.
such as training a forced aligner using children’s voices is [12] Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Y. Liu,
also part of our future work. It will help to generate more ‘‘Fastspeech 2: Fast and high-quality end-to-end text to speech,’’ 2020,
meaningful alignments for splitting the longer audio files arXiv:2006.04558.
[13] N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, ‘‘Neural speech synthesis
with transformer network,’’ Proc. AAAI Conf. Artif. Intell., vol. 33, no. 1,
13 [Link] Jul. 2019, pp. 6706–6713, doi: 10.1609/aaai.v33i01.33016706.

47640 VOLUME 10, 2022


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

[14] C. Miao, S. Liang, M. Chen, J. Ma, S. Wang, and J. Xiao, ‘‘Flow-TTS: [35] G. Yeung, R. Fan, and A. Alwan, ‘‘Fundamental frequency fea-
A non-autoregressive network for text to speech based on flow,’’ in Proc. ture normalization and data augmentation for child speech recog-
IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), May 2020, nition,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal Process.
pp. 7209–7213, doi: 10.1109/ICASSP40776.2020.9054484. (ICASSP2021), Toronto, ON, Canada, 2021, pp. 6993–6997, doi:
[15] J. Kim, S. Kim, J. Kong, and S. Yoon, ‘‘Glow-TTS: A generative 10.1109/ICASSP39728.2021.9413801.
flow for text-to-speech via monotonic alignment search,’’ Tech. Rep., [36] S. Shahnawazuddin, N. Adiga, H. K. Kathania, and B. T. Sai, ‘‘Creating
2020. speaker independent ASR system through prosody modification based data
[16] A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, augmentation,’’ Pattern Recognit. Lett., vol. 131, pp. 213–218, Mar. 2020,
A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, ‘‘WaveNet: doi: 10.1016/[Link].2019.12.019.
A generative model for raw audio,’’ 2016, arXiv:1609.03499. [37] S. Shahnawazuddin, R. Sinha, and G. Pradhan, ‘‘Pitch-normalized
[17] R. Prenger, R. Valle, and B. Catanzaro, ‘‘Waveglow: A flow-based gen- acoustic features for robust children’s speech recognition,’’ IEEE Sig-
erative network for speech synthesis,’’ in Proc. IEEE Int. Conf. Acoust., nal Process. Lett., vol. 24, no. 8, pp. 1128–1132, Aug. 2017, doi:
Speech Signal Process. (ICASSP), May 2019, pp. 3617–3621. 10.1109/LSP.2017.2705085.
[18] K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. sotelo, [38] S. Lee, A. Potamianos, and S. S. Narayanan, ‘‘Analysis of children’s
A. de Brébisson, Y. Bengio, and A. C. Courville, ‘‘MelGAN: Generative speech: Duration, pitch and formants,’’ in Proc. Eurospeech, 1997.
adversarial networks for conditional waveform synthesis,’’ in Proc. Adv. [39] S. Shahnawazuddin, N. Adiga, and H. K. Kathania, ‘‘Effect of prosody
Neural Inf. Process. Syst., vol. 32, 2019. modification on Children’s ASR,’’ IEEE Signal Process. Lett., vol. 24,
[19] J. Kong, J. Kim, and J. Bae, ‘‘HiFi-GAN: Generative adversarial networks no. 11, pp. 1749–1753, Nov. 2017, doi: 10.1109/LSP.2017.2756347.
for efficient and high-fidelity speech synthesis,’’ in Proc. Adv. Neural Inf. [40] M. Gerosa, D. Giuliani, S. Narayanan, and A. Potamianos, ‘‘A review of
Process. Syst., 2020, pp. 17022–17033. ASR technologies for children’s speech,’’ in Proc. 2nd Workshop Child,
[20] N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, Comput. Interact. (WOCCI), 2009, pp. 1–8.
E. Lockhart, F. Stimberg, A. Oord, S. Dieleman, and K. Kavukcuoglu, [41] G. Stemmer, C. Hacker, S. Steidl, and E. Nöth, ‘‘Acoustic normalization
‘‘Efficient neural audio synthesis,’’ in Proc. Int. Conf. Mach. Learn., 2018, of children’s speech,’’ in Proc. 8th Eur. Conf. Speech Commun. Technol.,
pp. 2410–2419. 2003.
[21] S. Arik, ‘‘Deep voice 2: Multi-speaker neural text-to-speech,’’ in Proc. [42] M. Gerosa, S. Lee, D. Giuliani, and S. Narayanan, ‘‘Analyzing children’s
31st Int. Conf. Neural Inf. Process. Syst. Red Hook, NY, USA, Curran speech: An acoustic study of consonants and consonant-vowel transition,’’
Associates Inc., 2017, pp. 2966–2974. in Proc. IEEE Int. Conf. Acoustics Speech Signal Process., 2006, pp. I–I,
[22] L. Wan, Q. Wang, A. Papir, and I. L. Moreno, ‘‘Generalized end-to-end loss doi: 10.1109/ICASSP.2006.1660040.
for speaker verification,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal [43] C. Li, Y. Qian, and M. Key, ‘‘Prosody usage optimization for children
Process. (ICASSP), Apr. 2018, pp. 4879–4883. speech recognition with zero resource children speech,’’ in Proc. Inter-
[23] G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, ‘‘End-to-end text- speech, 2019, pp. 3446–3450.
dependent speaker verification,’’ in Proc. IEEE Int. Conf. Acoustics, [44] S. Lee, A. Potamianos, and S. Narayanan, ‘‘Acoustics of children’s speech:
Speech Signal Process. (ICASSP), 2016, pp. 5115–5119. Developmental changes of temporal and spectral parameters,’’ J. Acoust.
[24] T.-H. Lo, F.-A. Chao, S.-Y. Weng, and B. Chen, ‘‘The NTNU system at the Soc. Amer., vol. 105, no. 3, pp. 1455–1468.
interspeech 2020 non-native children’s speech ASR challenge,’’ in Proc. [45] W. Ward, R. Cole, and S. Pradhan, ‘‘My science tutor and the MyST
Interspeech, Oct. 2020, pp. 1–5. corpus,’’ Tech. Rep., 2019.
[25] D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, [46] A. Nagrani, J. S. Chung, and A. Zisserman, ‘‘VoxCeleb: A large-
‘‘X-vectors: Robust DNN embeddings for speaker recognition,’’ in scale speaker identification dataset,’’ in Proc. Interspeech, Aug. 2017,
Proc. IEEE Int. Conf. Acoust., Speech Signal Process., Apr. 2018, pp. 1–5.
pp. 5329–5333, doi: 10.1109/ICASSP.2018.8461375. [47] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, ‘‘Librispeech: An
[26] E. Cooper, C. I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and ASR corpus based on public domain audio books,’’ in Proc. IEEE Int. Conf.
J. Yamagishi, ‘‘Zero-shot multi-speaker text-to-speech with state-of-the- Acoust., Speech Signal Process. (ICASSP), Apr. 2015, pp. 5206–5210, doi:
art neural speaker embeddings,’’ in Proc. IEEE Int. Conf. Acoust., Speech 10.1109/ICASSP.2015.7178964.
Signal Process. (ICASSP), May 2020, pp. 6184–6188. [48] C. Veaux, J. Yamagishi, and K. MacDonald, ‘‘CSTR VCTK corpus:
[27] M. Chen, X. Tan, Y. Ren, J. Xu, H. Sun, S. Zhao, and T. Qin, ‘‘MultiSpeech: English multi-speaker corpus for CSTR voice cloning toolkit,’’ Centre
Multi-speaker text to speech with transformer,’’ in Proc. Interspeech, Speech Technol. Res. (CSTR), Univ. Edinburgh, Tech. Rep., 2019, doi:
Oct. 2020, pp. 4024–4028. 10.7488/ds/2645.
[28] R. Valle, J. Li, R. Prenger, and B. Catanzaro, ‘‘Mellotron: Multispeaker [49] H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and
expressive voice synthesis by conditioning on rhythm, pitch and global Y. Wu, ‘‘LibriTTS: A corpus derived from LibriSpeech for text-to-speech,’’
style tokens,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. 2019, arXiv:1904.02882.
(ICASSP), May 2020, pp. 6189–6193. [50] The LJ Speech Dataset. Accessed: Mar. 15, 2021. [Online]. Available:
[29] A. Kulkarni, V. Colotte, and D. Jouvet, ‘‘Improving latent representation [Link]
for end-to-end multispeaker expressive text to speech system,’’ Tech. [51] CorentinJ/Real-Time-Voice-Cloning: Clone a Voice in 5 Seconds to Gen-
Rep. ffhal-02978485v1f, 2020. erate Arbitrary Speech in Real-Time. Accessed: May 27, 2021. [Online].
[30] E. Cooper, C.-I. Lai, Y. Yasuda, and J. Yamagishi, ‘‘Can speaker aug- Available: [Link]
mentation improve multi-speaker end-to-end TTS?’’ in Proc. Interspeech, [52] L. Mcinnes and J. Healy, ‘‘UMAP: Uniform manifold approximation and
Oct. 2020, pp. 1–5. projection for dimension reduction,’’ 2018, arXiv:1802.03426.
[31] M. Chen, M. Chen, S. Liang, J. Ma, L. Chen, S. Wang, and J. Xiao, ‘‘Cross- [53] Fatchord/WaveRNN: WaveRNN Vocoder + TTS. Accessed: May 27, 2021.
lingual, multi-speaker text-to-speech synthesis using neural speaker [Online]. Available: [Link]
embedding,’’ in Proc. Interspeech, Sep. 2019, pp. 2105–2109, doi: [54] P.-C. Hsu, C.-H. Wang, A. T. Liu, and H.-Y. Lee, ‘‘Towards robust neural
10.21437/Interspeech.2019-1632. vocoding for speech generation: A survey,’’ 2019, arXiv:1912.02461.
[32] W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, [55] J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz,
J. Raiman, and J. Miller, ‘‘Deep voice 3: Scaling text-to-speech with convo- R. Barra-Chicote, A. Moinet, and V. Aggarwal, ‘‘Towards achieving robust
lutional sequence learning,’’ Univ. California, Berkeley, Tech. Rep., 2017. universal neural vocoding,’’ 2018, arXiv:1811.06292.
[33] Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, [56] P. L. Tobing and T. Toda, ‘‘High-fidelity and low-latency universal neural
I. L. Moreno, and Y. Wu, ‘‘Transfer learning from speaker verification to vocoder based on multiband WaveRNN with data-driven linear prediction
multispeaker text-to-speech synthesis,’’ in Proc. Adv. Neural Inf. Process. for discrete waveform modeling,’’ in Proc. Interspeech, 2021, p. 2105.
Syst., Dec. 2018, pp. 4480–4490. [57] D. Paul, Y. Pantazis, and Y. Stylianou, ‘‘Speaker conditional WaveRNN:
[34] R. J. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, Towards universal neural vocoder for unseen speaker and recording condi-
J. Shor, R. Weiss, R. Clark, and R. A. Saurous, ‘‘Towards end-to- tions,’’ in Proc. Interspeech, 2020.
end prosody transfer for expressive speech synthesis with tacotron,’’ in [58] Z. Cai, C. Zhang, and M. Li, ‘‘From speaker verification to multispeaker
Proc. 35th Int. Conf. Mach. Learn., vol. 80. Stockholm, Sweden, 2018, speech synthesis, deep transfer with feedback constraint,’’ in Proc. Inter-
pp. 4693–4702. speech, 2020, pp. 3974–3978, doi: 10.21437/Interspeech.2020-1032.

VOLUME 10, 2022 47641


R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results

[59] B. Naderi and R. Cutler, ‘‘An open source implementation of ITU- DAN BIGIOI (Graduate Student Member, IEEE)
T recommendation P.808 with validation,’’ in Proc. Interspeech, 2020, received the bachelor’s degree in electronic and
pp. 2862–2866, doi: 10.21437/Interspeech.2020-2665. computer engineering from the National Univer-
[60] E. H. Rothauser, ‘‘IEEE recommended practice for speech quality mea- sity of Ireland Galway, in 2020. Upon graduat-
surements,’’ IEEE Trans. Audio Electroacoustics, vol. AU-17, no. 3, ing, he worked as a Research Assistant at NUIG
pp. 225–246, Jun. 1969. studying the text-to-speech and speaker recog-
[61] M. Viswanathan and M. Viswanathan, ‘‘Measuring speech quality for text-
to-speech systems: Development and assessment of a modified mean opin- nition methods under the DAVID (Data-Center
ion score (MOS) scale,’’ Comput. Speech Lang., vol. 19, no. 1, pp. 55–83, Audio/Visual Intelligence on-Device) Project.
Jan. 2005, doi: 10.1016/[Link].2003.12.001. Currently, he is working on his Ph.D. at NUIG,
[62] M. Chinen, F. S. C. Lim, J. Skoglund, N. Gureev, and A. Hines, ‘‘ViSQOL sponsored by D-REAL and the SFI Centre for
v3: An open source production ready objective speech and audio metric,’’ Research Training in Digitally Enhanced Reality. His research interests
Tech. Rep., Oct. 2021. include novel deep learning-based techniques for automatic speech dubbing
[63] C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, ‘‘A short- and discovering new ways to process multimodal audio/visual data.
time objective intelligibility measure for time-frequency weighted noisy
speech,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal Process.,
Mar. 2010, pp. 4214–4217, doi: 10.1109/ICASSP.2010.5495701.
[64] T. H. Falk, C. Zheng, and W.-Y. Chan, ‘‘A non-intrusive quality and
intelligibility measure of reverberant and dereverberated speech,’’ IEEE
Trans. Audio, Speech, Language Process., vol. 18, no. 7, pp. 1766–1774,
Sep. 2010, doi: 10.1109/TASL.2010.2052247.
[65] A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, ‘‘Percep-
tual evaluation of speech quality (PESQ)—A new method for speech
quality assessment of telephone networks and codecs,’’ in Proc. IEEE
Int. Conf. Acoust., Speech, Signal Process., May 2001, pp. 749–752, doi:
10.1109/ICASSP.2001.941023. PETER CORCORAN (Fellow, IEEE) is currently
[66] C.-C. Lo, S. W. Fu, W. C. Huang, X. Wang, J. Yamagishi, Y. Tsao, holding the Personal Chair of Electronic Engineer-
and H. M. Wang, ‘‘MOSNet: Deep learning-based objective assessment ing with the College of Science and Engineering,
for voice conversion,’’ in Proc. Interspeech, 2019, pp. 1541–1545, doi: National University of Ireland Galway (NUIG).
10.21437/Interspeech.2019-2003. He was the Co-Founder of several start-up com-
[67] M. Schoeffler, S. Bartoschek, F.-R. Stöter, M. Roess, S. Westphal, B. Edler,
panies, notably FotoNation (currently the Imaging
and J. Herre, ‘‘WebMUSHRA—A comprehensive framework for web-
based listening tests,’’ J. Open Res. Softw., vol. 6, no. 1, p. 8, Feb. 2018, Division, Xperi Corporation). He has more than
doi: 10.5334/jors.187. 600 cited technical publications and patents, more
[68] F. Ribeiro, D. Florencio, C. Zhang, and M. Seltzer, ‘‘CROWDMOS: than 120 peer-reviewed journal articles, 160 inter-
An approach for crowdsourcing mean opinion score studies,’’ in Proc. national conference papers, and a co-inventor on
IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), May 2011, more than 300 granted U.S. patents. He is an IEEE Fellow recognized for
pp. 2416–2419. his contributions to digital camera technologies, notably in-camera red-eye
[69] S. Zielinski, P. Hardisty, C. Hummersone, and F. Rumsey, ‘‘Potential biases correction and facial detection. He is also a member of the IEEE Consumer
in MUSHRA listening tests,’’ in Proc. Audio Eng. Soc. 123rd Audio Eng. Technology Society for more than 25 years and the Founding Editor of IEEE
Soc. Conv., vol. 2, 2007, pp. 1–10. Consumer Electronics Magazine.
[70] M. Wester, C. Valentini-Botinhao, and G. E. Henter, ‘‘Are we using
enough listeners? No! An empirically-supported critique of Interspeech
2014 TTS evaluations,’’ in Proc. Interspeech, 2015, pp, 3476–3480, doi:
10.21437/Interspeech.2015-689.

RISHABH JAIN (Graduate Student Member,


IEEE) received the [Link]. degree in computer
science and engineering from the Vellore Insti-
tute of Technology (VIT), in 2019, and the M.S.
degree in data analytics from the National Univer-
sity of Ireland Galway (NUIG), in 2020, where HORIA CUCU (Member, IEEE) received the B.S.
he is currently pursuing the Ph.D. degree. He is and M.S. degrees in applied electronics and the
also working as a Research Assistant at NUIG Ph.D. degree in electronics and telecommunica-
under DAVID (Data-Center Audio/Visual Intelli- tion engineering from the University Politehnica
gence on-Device) Project. His research interests of Bucharest (UPB), Romania, in 2008 and 2011,
include machine learning and artificial intelligence specifically in domain respectively.
of speech understanding, text-to-speech, speaker recognition, and automatic From 2010 to 2017, he was a Teaching Assistant
speech recognition. and then a Lecturer at UPB, where he is cur-
rently working as an Associate Professor. In this
MARIAM YAHAYAH YIWERE received the position, he authored over 75 scientific papers
Bachelor of Science degree from the Department in international conferences and journals, served as the project director
of Computer Science, Kwame Nkrumah Univer- for seven research projects, and contributed as a researcher to ten other
sity of Science and Technology, Kumasi, Ghana, research grants. He holds two patents. In addition, he founded and leads
in 2012, and the Master of Engineering and Ph.D. Zevo Technology, a speech start-up dedicated to integrating state-of-the-
degrees from the Department of Computer Engi- art speech technologies in various commercial applications. His research
neering, Hanbat National University, South Korea, interests include machine/deep learning and artificial intelligence, with a
in August 2015 and February 2020, respectively. special focus on automatic speech and speaker recognition, text-to-speech
Since October 2020, she has been working on the synthesis, and speech emotion recognition.
DTIF/DAVID Project as a Postdoctoral Researcher Dr. Cucu was awarded the Romanian Academy Prize ‘‘Mihail Drăgă-
with the College of Science and Engineering, National University of Ireland nescu’’ (2016) for Outstanding Research Contributions in Spoken Language
Galway, Galway. Her research interests include text-to-speech synthesis, Technology, after developing the first large-vocabulary automatic speech
speaker recognition and verification, sound source localization, deep learn- recognition system for the Romanian language.
ing, and computer vision.
47642 VOLUME 10, 2022

You might also like