Child Speech Synthesis TTS Pipeline Evaluation
Child Speech Synthesis TTS Pipeline Evaluation
May 6, 2022.
Digital Object Identifier 10.1109/ACCESS.2022.3170836
ABSTRACT Speech synthesis has come a long way as current text-to-speech (TTS) models can now generate
natural human-sounding speech. However, most of the TTS research focuses on using adult speech data
and there has been very limited work done on child speech synthesis. This study developed and validated a
training pipeline for fine-tuning state-of-the-art (SOTA) neural TTS models using child speech datasets. This
approach adopts a multi-speaker TTS retuning workflow to provide a transfer-learning pipeline. A publicly
available child speech dataset was cleaned to provide a smaller subset of approximately 19 hours, which
formed the basis of our fine-tuning experiments. Both subjective and objective evaluations were performed
using a pretrained MOSNet for objective evaluation and a novel subjective framework for mean opinion
score (MOS) evaluations. Subjective evaluations achieved the MOS of 3.95 for speech intelligibility, 3.89 for
voice naturalness, and 3.96 for voice consistency. Objective evaluation using a pretrained MOSNet showed a
strong correlation between real and synthetic child voices. Speaker similarity was also verified by calculating
the cosine similarity between the embeddings of utterances. An automatic speech recognition (ASR) model
is also used to provide a word error rate (WER) comparison between the real and synthetic child voices. The
final trained TTS model was able to synthesize child-like speech from reference audio samples as short as
5 seconds.
INDEX TERMS Text-to-speech, child speech synthesis, tacotron, multi-speaker TTS, alternative WaveRNN,
MOSNet, subjective MOS.
I. INTRODUCTION voice services, TTS models are also important, and the most
The bulk of recent research into human speech has focused on advanced models can incorporate emotional and prosodic
neural network techniques to improve speech understanding elements into the generated speech output.
and recognition or to provide simplified, high-quality text- More recent research into low-resource languages and
to-speech (TTS) models that can directly convert written other low-resource aspects of human speech, such as accented
text into natural speech. The most highly developed domain and prosody-aligned speech has started to see improvements
for such research has a focus on spoken English and is for both ASR and TTS [1]. Another aspect of human speech
based on native-speaker adult voice data samples. Automated of growing importance is that of child speech. Child speech
speech recognition (ASR) is a core element of modern con- differs significantly from those of adult speech, falling into a
sumer technology user interfaces employed in smart-speaker narrow range of variation and with higher pitch levels. Fur-
and voice command interfaces. For interactive chatbot and thermore, children’s speech patterns are more inarticulate and
can vary widely in terms of volume, pacing, and emotional
expressivity. These challenges are further amplified by the
The associate editor coordinating the review of this manuscript and relatively small number of public child speech corpora that
approving it for publication was Juan Wang . are available with useful annotations.
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
47628 VOLUME 10, 2022
R. Jain et al.: TTS Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results
Current work done on TTS for child’s voices is limited. trend of data-hungry DNN-based TTS, TTS for children has
This is mainly due to the lack of child voice datasets and dif- practically been neglected due to the lack of large publicly
ficulty in creating such datasets. As TTS models require hun- available children’s speech datasets suitable for training such
dreds of hours of annotated data for training [2], performing networks. Prior to this DNN era, researchers worked on TTS
TTS for child voices can be quite challenging. The focus of for children using HMM-based models [3], [6].
this work is to explore the potential of state-of-the-art (SOTA) Collecting data for child speech research can be a chal-
TTS to build a pipeline for the synthesis of children’s voices lenging task. Most TTS datasets are created in studios with
with low data requirements. More specifically, if we can build expensive equipment: an adult will be using a microphone to
such a pipeline and demonstrate that it can reliably synthesize create a clean, noiseless, easy to understand, and meaningful
a useful number of distinct children’s voices, this pipeline audio. This task is not easy to produce and even more difficult
would enable the creation of large synthetic datasets that to implement with a child.
could further improve other aspects of child speech research One of the main differences between adult speech and child
such as automatic speech recognition (ASR), speaker recog- speech is the fundamental frequency. The pitch for children is
nition, etc. To better elaborate on this hypothesis, it is useful significantly higher than that of an adult [35]–[38]. The pitch
to review current SOTA in TTS technologies, followed by a for an adult voice lies between 70 to 250 Hz whereas the pitch
similar consideration for review in child speech research. for the children’s speech is between 200 to 500 Hz [39]. There
is also a difference in the speaking rate of children. It was
A. RELATED RESEARCH IN TTS noticed that average phoneme duration is longer in children,
Early research work on TTS synthesis can be traced therefore, leading to longer speaking rates as compared to
back to four/five decades ago when the task of TTS adult speech [38], [40]–[42]. The vocal tract of an adult is
was commonly tackled using concatenative and parametric larger as compared to children’s vocal tract and therefore
approaches [3]–[7]. Although these early methods were suc- produces different prosody features as compared to an adult
cessful in generating speech from text, they generally lacked voice [43], [44]. Hence, a substantial difference in children’s
naturalness. The audio generated using these approaches was voice characteristics and features can be seen as compared to
kind of muffled and sounded very robotic. an adult voice.
Recent state-of-the-art TTS models are largely based on Our work aims to solve the problem of TTS for children
deep neural networks (DNN) and can achieve more natural- using DNNs. To solve this problem, the huge challenge of
sounding/human-like synthesized speech. With the introduc- limited publicly available children’s speech datasets must
tion of Tacotron [8], a neural sequence to sequence the TTS first be overcome. To this end, this study considered the use
model, the quality of speech synthesis improved significantly. of an existing multi-speaker children’s speech dataset [45],
While there are newer approaches that are more efficient or which comes with an incomplete set of utterance transcrip-
use smaller models, etc., it is still representative of SOTA tions. In addition, this dataset has a lot of unusable data,
for the quality of the synthesized speech and is used as a such as empty/blank entries, extremely long entries as well as
benchmark for comparison with newer methods. Nonethe- inaccurate transcriptions. Firstly, the dataset is cleaned up to
less, Tacotron TTS is not very robust as it sometimes skips create a subset that is suitable for training a neural TTS model.
certain words and it also suffers from low inference speed Secondly, with the cleaned-up dataset, a multi-speaker TTS
[9]. Several methods have since been proposed to improve model is trained to generate synthetic speech for multiple
upon it such as Tacotron2 [10], FastSpeech [11], FastSpeech2 child speakers as a proof of concept for children’s TTS. The
[12], Transformer TTS [13], FlowTTS [14], GlowTTS [15], training involved fine-tuning an existing adult multi-speaker
etc. Similarly, there have been several improvements over TTS model [33] by way of transfer learning, with a few
the quality of synthesized waveforms by the introduction modifications as explained in later sections. This approach
of SOTA Vocoders such as WaveNet [16], WaveGlow [17], involves the training of a separate speaker verification model,
MelGAN [18], Hifi-Gan [19], WaveRNN [20], etc. These and it was preferred because it reduces the problem at hand
TTS models supported single speaker synthesis, but Deep- in two ways:
voice2 [21], introduced the use of speaker verification mod- 1) To train the speaker verification network, transcrip-
els [22]–[25] to achieve Multi-speaker TTS [26]–[34]. tions for the speech dataset are not required. Only the
speaker identities for the utterances are needed and it
B. CHILD SPEECH – LITERATURE AND CHALLENGES can also be trained on noisy speech without any nega-
While all SOTA TTS systems rely on large datasets to train, tive effects. This means that even the noisy children’s
the datasets mostly comprise speech taken from adult native speech dataset, which has incomplete transcriptions,
English speakers; hence, for low-resource languages and can be useful in training the verification model.
other target groups such as non-native adult speakers and 2) Being a transfer learning process, the pretrained TTS
child speakers, there remain challenges developing effective model can be finetuned sufficiently using the resulting
and suitable TTS models. Specifically, in comparison with cleaned set of children’s speech data.
adult TTS, child TTS has gained very little to no atten- Subjective and Objective Evaluation performed on the syn-
tion from the TTS research community. With the current thesized child voices confirms that the child voices generated
TABLE 1. Dataset used in this work. TABLE 2. MyST dataset comparison [complete vs with transcript].
not very accurate for the child speech and there were a
lot of mismatches between the transcripts and audio files.
This was probably due to fact that the pretrained forced
aligner was trained on adult speech and doesn’t work very
well for aligning child speech. Therefore, TinyMyST was
• Transcription containing text with no phonetic used (as described earlier) for performing all the child TTS
information. experiments.
• ‘‘(()) (()) (())’’
• Repetition of words/stammering noticed in children’s 3) DATA PREPROCESSING FOR TTS USAGE
voices. LibriSpeech and TinyMyST datasets were preprocessed as
• ‘‘um we measured how big a millimeter meter is a per the guidelines mentioned in LibriTTS [49]. The LibriTTS
meter and a kolome- a ∗ kilometer ∗’’ dataset was specifically created for TTS research, therefore
Our examination of MyST led us to further clean the MyST similar guidelines were followed in our experiments. The
dataset for TTS training. In this process a subset of MyST, following changes were made:
hereafter referred to as TinyMyST was created. • Audio files were converted to 16-bit depth audio files
with 24Khz sampling rate (WAV format), This was done
2) TINYMYST using the pydub2 audio library.
It is a small subset of the MyST dataset created using var- • Text data was normalized by replacing abbreviations and
ious pre-processing scripts to make the data suitable for punctuations.
TTS acoustic model training. MyST was cleaned to select • Whitespaces were normalized
only audio files with existing transcriptions. All audio files • All characters were made uppercase.
lesser than 10 seconds and greater than 15 seconds were
removed. The utterances shorter than 10 seconds contained B. MULTI-SPEAKER TTS MODEL
mostly noise or unintelligible speech and those longer than
The neural network used to achieve TTS for children in this
15 seconds were removed to avoid GPU memory overflow
study is based on [33], It works by combining a speaker verifi-
during training. All the transcript files were converted from
cation network with the SOTA Tacotron TTS model. Though
.trn format to .txt file format.
Tacotron is SOTA for TTS, it was designed to be trained using
The TinyMyST dataset still contains a lot of noisy data.
a single-speaker speech dataset such as the LJSpeech [50]
Some of the excessively noisy data were removed man-
dataset, hence, it can only synthesize speech with acoustic
ually by inspecting the transcripts and listening to the
characteristics of the single speaker whose data was used
audio samples. The data obtained after cleaning contained
in training. To function effectively for multiple speakers,
7152 utterances and accounted for 19.22 hours. A detailed
Tacotron needs to be adapted for that purpose. This adaptation
comparison of MyST and TinyMyST datasets was performed
has been achieved in this multi-speaker TTS model [33] by
to see differences in the two datasets in terms of speaker iden-
introducing different speaker identities in the form of speaker
tities and utterances (see Table 3). TinyMyST dataset on aver-
embeddings as additional input to the Tacotron network. As a
age contained 1.72 minutes per speaker having 670 speakers.
result, the multi-speaker TTS [33] comprises three different
Speaker identity ‘013023’ had the most data with 8.77 min-
neural network models, each of which focuses on a specific
utes and speaker identity ‘018216’ has the least data with
subtask namely, Speaker Encoder used for speaker verifica-
10.01 seconds. The speaker ‘013023’ had the most data in
tion task, Acoustic model used for spectrogram synthesis,
MyST as well to be around 110 minutes.
and a Vocoder for audio waveform generation (as shown in
To extract more TTS usable data, an audio sample from
Figure 1).
more than 15 seconds long can be used to split them into
For our work, generalized end-to-end (GE2E) loss was
smaller chunks. A forced aligner1 is used to align the audio
used for speaker verification [22], Tacotron1 as an acoustic
files with transcripts. Time alignment information from the
model [8], and WaveRNN as Vocoder [20]. The original
alignments to split the longer audio files into smaller sam-
approach [33] is adapted for child speech synthesis by first
ples, however, it was observed that the audio alignment was
1 [Link] 2 [Link]
pretraining the model on an adult speech dataset after which, TABLE 4. Speaker encoder training details.
it is fine-tuned with the child speech dataset.
The speaker encoder generates speaker embeddings,
encoding speaker identity information extracted from the
utterances. Similar voices are mapped closer to each other
in a latent space representation. The acoustic model gen-
erates spectrograms from text conditioned on the speaker
embeddings. The vocoder then converts these spectrograms
into audio waveforms. At inference time, a short refer-
ence utterance (ground truth) of a child’s voice is passed
through the speaker encoder to generate the corresponding
speaker embeddings, on which the acoustic model will be
conditioned. The three different neural network models are
described as follows.
1) SPEAKER ENCODER
The first stage of the multi-speaker TTS training involves the
training of a speaker verification (speaker encoder) model.
Speaker Verification is the process of determining if an utter-
ance belongs to a specific speaker. The speaker encoder is
used to train the model for the speaker verification task using FIGURE 2. Pipeline for Speaker Encoder training. The dotted line
represents the training loop for the Speaker Encoder training.
a mix of noisy and clean speech data without transcripts.
The data used consists of both adult and child speech data
from thousands of speakers (see Table 4). This was done to relatively insignificant improvements were seen in the EER
introduce both child and adult speakers in the model for better after this point.
generalization. The output of this model conditions the acous- All the datasets were pre-processed into the coding format
tic model to generate the required mel-spectrograms from required for training the encoder as described in [51]. Even
a reference speech signal of the target speaker. The model though half the MyST dataset is not transcribed, the complete
is trained to capture the characteristic features of different MyST dataset can be used for Speaker Encoder training as it
speakers. does not require any transcription data. The pipeline for the
The model takes input as log mel-spectrograms computed speaker encoder training can be seen in Figure 2.
from utterances of each speaker, trains using the GE2E loss A UMAP projection [52] is created to visualize the training
and converts them into a fixed dimensional vector called by taking a random set of 10 utterances from 10 speakers.
d-vectors. These d-vectors are optimized over GE2E loss Utterances with similar embeddings are located close to each
to differentiate the speakers, such that the same speakers other in the latent space representation and have similar
have embeddings with high cosine similarity and different speaker characteristics.
speakers are far apart in the embedding space. This model creates individual clusters of speaker embed-
During training, complete utterances are segmented into dings as can be seen in the UMAP projection (see Figure 3.
partial utterances of 1.6 seconds. These parameters were kept Each point on UMAP represents an utterance. The same color
the same as explained by authors [51], [22]. The utterance points represent the same speaker. Encoder gradually learns
embedding is calculated using 800ms windows for inference, to separate the speakers. Initially, there is a lot of overlap
with a 50% overlap. The silence was removed from the utter- across speakers, but eventually, each speaker has their utter-
ances using the webrtcvad3 tool for Voice Activity Detection ances clustered and well separated from the other speakers.
(VAD). Each segment is passed through the network individ- The training evolves with increased training steps.
ually, the outputs are averaged and normalized to create the
final utterance embedding as described in [22]. 2) TACOTRON ACOUSTIC MODEL
The encoder model is trained using 4 datasets, MyST, Vox- For the speech spectrogram synthesis, the TTS model archi-
Celeb1, LibriSpeech, and VCTK. Equal Error Rate (EER) is tecture and hyperparameters used in this study are the
used as a metric for the validation of the speaker encoder. The same as in the work of [51] (More details are provided
default EER metric from [51] is used in this work as authors in Section III). The authors used a modified version of the
of [33] have not explicitly specified the training, test, and original Tacotron architecture [8]. The model consists of an
validation criterion they are using for EER calculation. The encoder, an attention-based decoder, and a post-processing
EER values are presented in Table 4. The model trained for network. Since Tacotron is originally a single-speaker TTS
one million steps was used in the multi-speaker TTS model as model, it was modified to work for multi-speaker TTS by con-
necting the speaker encoder to it. Speaker embeddings from
3 [Link] the encoder are concatenated with text (character/phoneme)
FIGURE 6. Alignment plot for Tacotron 2 trained with MyST dataset for up
to 200k steps.
FIGURE 5. Pipeline for Vocoder training. Models with solid contours are
pretrained models. The dotted contour represents the training loop for
Vocoder.
in a noise-free environment. They were asked to rate each TABLE 5. MOS (from 1 to 5) explained for speech intelligibility and voice
naturalness.
of the synthetic voices assigned to them from a range of
1 to 5, for each of the different categories in two phases. The
categories included Speech Intelligibility, Voice Naturalness,
and Voice Consistency. Voice Consistency contained three
sub-categories of its own namely, Start of Phrase quality,
Middle of Phrase quality, and End of Phrase Quality.
Evaluation data was provided in a OneDrive Environment.
All the synthetic voices were shared in a common OneDrive
folder to the evaluators and a common spreadsheet was circu-
lated containing the utterance ID of Harvard sentences used
for synthesizing a child’s voice. While listening to many
different natural child voices, it was also noticed that recorded
child audio can be a difficult task to understand if not pro-
vided with a suitable transcript. Some of the child’s speech
can be non-meaningful as mentioned in problems with the
MyST section. After performing many different tests and
trials using child speech, the use of transcripts as a part of TABLE 6. MOS from phase-I evaluation with 95% confidence interval
MOS-based evaluation is considered to be a more natural
way of evaluating child speech. Therefore, corresponding
transcript information is also provided in the spreadsheet to
each evaluator to base their conclusion on ‘what they hear
in child audio’ and ‘what they read in child transcripts’.
This way more coherent patterns can be observed among the The spreadsheet was later analyzed to get the final mean
phonemes and graphemes in a child’s voice for each of the opinion score in each category. MOS of 3.88 for voice nat-
mentioned categories. An example of this spreadsheet can be uralness and 4.13 for speech intelligibility was observed as
seen here.9 seen in Table 6.
By performing the evaluation using OneDrive environ- An average score for each of the 720 sentences was cal-
ment, it was easy to distribute the synthetic speech files to dif- culated for the combined value of speech intelligibility and
ferent evaluators without having to spend time and resources voice naturalness. All the 720 sentences were sorted into
on expensive Mushra-based evaluations [67] or crowdsourc- difficult and easy sentences with respect to the children’s
ing the evaluation task on platforms like Amazon Mechani- linguistic capabilities. This was done to keep track of Harvard
cal Turk (AMT) [59], [68]. Mushra-based evaluations were sentences where synthesized speech becomes unintelligible
also avoided due to potential biases that can occur in these and inarticulate.
tests and how these biases can impact synthetic child voice
evaluation for MOS [69]. Most of these TTS evaluations 2) PHASE-II EVALUATION
have been conducted before with synthetic adult speech, this After our phase-I evaluation, a common set of sentences
novel synthetic evaluation is implemented for first-time with were observed where pronunciation sounds unintelligible
synthetic child speech. Using a common spreadsheet made it at the start, middle, or end of sentences for specific
effective to perform analysis of spreadsheet for MOS using words/phonemes. There was an inconsistency in voice qual-
pandas and other python-based tools. ity. These sets of sentences are the ones that were not learned
properly during training or were missing in the training
1) PHASE-I EVALUATION dataset for child audio. To make a note of these sentences,
For the phase-I evaluation, all 720 Harvard sentences were extra categories of ‘Voice Consistency’ were added to the
generated using our proposed TTS method. Two random ref- phase-I evaluation. Therefore, all the 3 sub-categories under
erence utterances were selected from the TinyMyST dataset Voice Consistency were used in the second phase of the eval-
and were used to generate all the Harvard sentences. These uation. These subcategories included ‘Start of Phrase Qual-
720 sentences were shared among 5 evaluators in a spread- ity’, ‘Middle of phrase quality’ and ‘End of phrase quality’.
sheet document, who rated the voices from 1 to 5 based The MOS ratings from 1 to 5 for each of these categories were
on Speech Intelligibility and Voice Naturalness. The MOS also explained in the spreadsheet as mentioned in Table 7.
ratings from 1 to 5 were further explained in the spreadsheet For the second phase of evaluation, the evaluation was
file as can be seen in Table 5. undertaken by 20 evaluators divided into 4 groups. This was
done as per the guidelines mentioned in [70] for performing
MOS evaluations. For each group, a speaker identity was
9 Example Spreadsheet: [Link] selected from the TinyMyST dataset. All the speaker iden-
blob/main/synthetic%20evaluation%20example/[Link] tities were sorted, and the top 20 speaker identities were
TABLE 7. MOS (from 1 to 5) explained for voice consistency and its three TABLE 9. MOS from phase-II evaluation with 95% confidence interval.
sub-categories.
TABLE 10. MOS natural speech VS MOS synthetic speech with 95% TABLE 11. MOSNet output for 5 samples.
confidence interval.
TABLE 12. MOSNet output for 50 samples with 95% confidence interval.
TABLE 13. WER on adult speech, real child speech and synthetic child to increase the training dataset. The information collected
speech.
from our subjective evaluation such as voice consistency in
Harvard sentences will be used to improve child speech.
This information will be used to collect better TTS-based
child speech data based on Harvard sentences to accord with
‘end of the phrase’ information loss and voice inconsistency
75% of the synthetic speech with a relative difference of observed with our current results. The use of synthetically
10 WER when compared with real child speech recognized generated child speech to improve other areas of child speech
by the same model for the same speakers. research such as ASR and speaker recognition will also be
investigated in future work. TTS-generated child voices can
V. CONCLUSION AND FUTURE WORK be used as a data augmentation technique for training these
In this paper, a pipeline for generating synthetic child speech models with additional data. It is also intended to use the
in a limited training data scenario is proposed. A small subjective evaluation method proposed in this paper for per-
set of child speech data is created by cleaning an existing forming all future subjective evaluations with TTS generated
child speech dataset and making it suitable for TTS training. child speech.
A transfer learning approach is used to train our model with
adult speech data in a pretraining setting and child speech data ACKNOWLEDGMENT
as low as 19 hours for fine-tuning. MOSNet based objective The authors would like to thank experts from Xperi-Ireland:
evaluation shows a high correlation between real and synthe- Gabriel Costache, George Sterpu, and the rest of the
sized child voices. A subjective evaluation method suitable team members for providing their expertise and feedback
for synthesized child speech is also proposed and demon- throughout.
strated. Subjective MOS of synthesized voices is observed
as 3.95 for speech intelligibility, 3.89 for voice naturalness, REFERENCES
and 3.96 for overall voice consistency which is very close [1] J. Xu, X. Tan, Y. Ren, T. Qin, J. Li, S. Zhao, and T. Y. Liu, ‘‘LRSpeech:
Extremely low-resource speech synthesis and recognition,’’ in Proc. 26th
to Natural speech MOS. These MOS values tell us about ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, Assoc. Comput.
how good the synthesized child voices are. However, voice Mach. New York, NY, USA, 2020, pp. 2802–2812, doi: 10.1145/3394486.
inconsistency for ‘End of phrase quality’ containing noise and [2] K. R. Prajwal and C V Jawahar, ‘‘Data-efficient training strategies for
neural TTS systems,’’ in Proc. 8th ACM IKDD CODS, 26th COMAD.
unintelligible information was also observed. There is scope New York, NY, USA: Association for Computing Machinery, 2021,
for improvement for these phrases. WER for synthetic child pp. 223–227, doi: 10.1145/3430984.3431034.
voices using a pretrained adult speech wav2vec2 ASR model [3] O. Watts, J. Yamagishi, K. Berkling, and S. King, ‘‘HMM-based synthesis
of child speech,’’ in Proc. 1st Work. Child, Comput. Interact. (ICMI Post-
came to be 25.63 as compared to WER of real child voices Conf. Work., 2008.
of 15.27. Synthetic child speech samples can be viewed in [4] H. Zen, K. Tokuda, and A. W. Black, ‘‘Statistical parametric speech
our GitHub repository.13 Multi-speaker TTS can be the key synthesis,’’ Speech Commun., vol. 51, no. 11, pp. 1039–1064, Nov. 2009,
doi: 10.1016/[Link].2009.04.004.
to child speech synthesis with limited training data. Child [5] P. K. Muthukumar and A. W. Black, ‘‘A deep learning approach to
speakers with speech duration between 5-7 minutes in TTS data-driven parameterizations for statistical parametric speech synthesis,’’
training gave 81% average cosine similarity with a synthetic Carnegie Mellon University Pittsburgh, Pittsburgh, NA, USA, 2014.
[6] O. Watts, J. Yamagishi, S. King, and K. Berkling, ‘‘Synthesis of child
speech from the same speakers. This choice of the model speech with HMM adaptation and voice conversion,’’ IEEE Trans. Audio,
allows the TTS to learn useful speaker information which can Speech Language Process., vol. 18, no. 5, pp. 1005–1016, Jul. 2010, doi:
be leveraged to produce better quality synthetic voices even 10.1109/TASL.2009.2035029.
[7] R. Maia, H. Zen, and M. J. F. Gales, ‘‘Statistical parametric speech synthe-
with limited child speech. sis with joint estimation of acoustic and excitation model parameters,’’ in
For future work, our aim is to improve this method by Proc. SSW, 2010, pp. 88–93.
incorporating more information to our multi-speaker TTS [8] Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly,
Z. Yang, Y. Xiao, Z. Chen, S. Bengio, and Q. Le, ‘‘Tacotron: Towards end-
model such as duration predictor and energy as implemented to-end speech synthesis,’’ 2017, arxiv:1703.10135.
in FastSpeech2 [12]. The trained vocoder was also finetuned [9] J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu,
on the TinyMyST dataset. However, there was no significant ‘‘Non-attentive tacotron: Robust and controllable neural TTS synthesis
including unsupervised duration modeling,’’ 2020, arxiv:2010.04301.
improvement in the quality of the generated audio waveforms [10] J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen,
and an additional noise was observed in some of the synthesis. Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis,
More child speech data would be required to achieve any and Y. Wu, ‘‘Natural TTS synthesis by conditioning wavenet on MEL
spectrogram predictions,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal
significant improvement over the quality of the vocoder. It is Process. (ICASSP), Apr. 2018, pp. 2756–2761.
also intended to implement GAN-based SOTA Vocoders such [11] Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Y. Liu,
as HiFi-GAN [19] for future experiments. More experiments ‘‘FastSpeech: Fast, robust and controllable text to speech,’’ in Proc. Adv.
Neural Inf. Process. Syst., 2019, vol. 32.
such as training a forced aligner using children’s voices is [12] Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Y. Liu,
also part of our future work. It will help to generate more ‘‘Fastspeech 2: Fast and high-quality end-to-end text to speech,’’ 2020,
meaningful alignments for splitting the longer audio files arXiv:2006.04558.
[13] N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, ‘‘Neural speech synthesis
with transformer network,’’ Proc. AAAI Conf. Artif. Intell., vol. 33, no. 1,
13 [Link] Jul. 2019, pp. 6706–6713, doi: 10.1609/aaai.v33i01.33016706.
[14] C. Miao, S. Liang, M. Chen, J. Ma, S. Wang, and J. Xiao, ‘‘Flow-TTS: [35] G. Yeung, R. Fan, and A. Alwan, ‘‘Fundamental frequency fea-
A non-autoregressive network for text to speech based on flow,’’ in Proc. ture normalization and data augmentation for child speech recog-
IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), May 2020, nition,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal Process.
pp. 7209–7213, doi: 10.1109/ICASSP40776.2020.9054484. (ICASSP2021), Toronto, ON, Canada, 2021, pp. 6993–6997, doi:
[15] J. Kim, S. Kim, J. Kong, and S. Yoon, ‘‘Glow-TTS: A generative 10.1109/ICASSP39728.2021.9413801.
flow for text-to-speech via monotonic alignment search,’’ Tech. Rep., [36] S. Shahnawazuddin, N. Adiga, H. K. Kathania, and B. T. Sai, ‘‘Creating
2020. speaker independent ASR system through prosody modification based data
[16] A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, augmentation,’’ Pattern Recognit. Lett., vol. 131, pp. 213–218, Mar. 2020,
A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, ‘‘WaveNet: doi: 10.1016/[Link].2019.12.019.
A generative model for raw audio,’’ 2016, arXiv:1609.03499. [37] S. Shahnawazuddin, R. Sinha, and G. Pradhan, ‘‘Pitch-normalized
[17] R. Prenger, R. Valle, and B. Catanzaro, ‘‘Waveglow: A flow-based gen- acoustic features for robust children’s speech recognition,’’ IEEE Sig-
erative network for speech synthesis,’’ in Proc. IEEE Int. Conf. Acoust., nal Process. Lett., vol. 24, no. 8, pp. 1128–1132, Aug. 2017, doi:
Speech Signal Process. (ICASSP), May 2019, pp. 3617–3621. 10.1109/LSP.2017.2705085.
[18] K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. sotelo, [38] S. Lee, A. Potamianos, and S. S. Narayanan, ‘‘Analysis of children’s
A. de Brébisson, Y. Bengio, and A. C. Courville, ‘‘MelGAN: Generative speech: Duration, pitch and formants,’’ in Proc. Eurospeech, 1997.
adversarial networks for conditional waveform synthesis,’’ in Proc. Adv. [39] S. Shahnawazuddin, N. Adiga, and H. K. Kathania, ‘‘Effect of prosody
Neural Inf. Process. Syst., vol. 32, 2019. modification on Children’s ASR,’’ IEEE Signal Process. Lett., vol. 24,
[19] J. Kong, J. Kim, and J. Bae, ‘‘HiFi-GAN: Generative adversarial networks no. 11, pp. 1749–1753, Nov. 2017, doi: 10.1109/LSP.2017.2756347.
for efficient and high-fidelity speech synthesis,’’ in Proc. Adv. Neural Inf. [40] M. Gerosa, D. Giuliani, S. Narayanan, and A. Potamianos, ‘‘A review of
Process. Syst., 2020, pp. 17022–17033. ASR technologies for children’s speech,’’ in Proc. 2nd Workshop Child,
[20] N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, Comput. Interact. (WOCCI), 2009, pp. 1–8.
E. Lockhart, F. Stimberg, A. Oord, S. Dieleman, and K. Kavukcuoglu, [41] G. Stemmer, C. Hacker, S. Steidl, and E. Nöth, ‘‘Acoustic normalization
‘‘Efficient neural audio synthesis,’’ in Proc. Int. Conf. Mach. Learn., 2018, of children’s speech,’’ in Proc. 8th Eur. Conf. Speech Commun. Technol.,
pp. 2410–2419. 2003.
[21] S. Arik, ‘‘Deep voice 2: Multi-speaker neural text-to-speech,’’ in Proc. [42] M. Gerosa, S. Lee, D. Giuliani, and S. Narayanan, ‘‘Analyzing children’s
31st Int. Conf. Neural Inf. Process. Syst. Red Hook, NY, USA, Curran speech: An acoustic study of consonants and consonant-vowel transition,’’
Associates Inc., 2017, pp. 2966–2974. in Proc. IEEE Int. Conf. Acoustics Speech Signal Process., 2006, pp. I–I,
[22] L. Wan, Q. Wang, A. Papir, and I. L. Moreno, ‘‘Generalized end-to-end loss doi: 10.1109/ICASSP.2006.1660040.
for speaker verification,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal [43] C. Li, Y. Qian, and M. Key, ‘‘Prosody usage optimization for children
Process. (ICASSP), Apr. 2018, pp. 4879–4883. speech recognition with zero resource children speech,’’ in Proc. Inter-
[23] G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, ‘‘End-to-end text- speech, 2019, pp. 3446–3450.
dependent speaker verification,’’ in Proc. IEEE Int. Conf. Acoustics, [44] S. Lee, A. Potamianos, and S. Narayanan, ‘‘Acoustics of children’s speech:
Speech Signal Process. (ICASSP), 2016, pp. 5115–5119. Developmental changes of temporal and spectral parameters,’’ J. Acoust.
[24] T.-H. Lo, F.-A. Chao, S.-Y. Weng, and B. Chen, ‘‘The NTNU system at the Soc. Amer., vol. 105, no. 3, pp. 1455–1468.
interspeech 2020 non-native children’s speech ASR challenge,’’ in Proc. [45] W. Ward, R. Cole, and S. Pradhan, ‘‘My science tutor and the MyST
Interspeech, Oct. 2020, pp. 1–5. corpus,’’ Tech. Rep., 2019.
[25] D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, [46] A. Nagrani, J. S. Chung, and A. Zisserman, ‘‘VoxCeleb: A large-
‘‘X-vectors: Robust DNN embeddings for speaker recognition,’’ in scale speaker identification dataset,’’ in Proc. Interspeech, Aug. 2017,
Proc. IEEE Int. Conf. Acoust., Speech Signal Process., Apr. 2018, pp. 1–5.
pp. 5329–5333, doi: 10.1109/ICASSP.2018.8461375. [47] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, ‘‘Librispeech: An
[26] E. Cooper, C. I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and ASR corpus based on public domain audio books,’’ in Proc. IEEE Int. Conf.
J. Yamagishi, ‘‘Zero-shot multi-speaker text-to-speech with state-of-the- Acoust., Speech Signal Process. (ICASSP), Apr. 2015, pp. 5206–5210, doi:
art neural speaker embeddings,’’ in Proc. IEEE Int. Conf. Acoust., Speech 10.1109/ICASSP.2015.7178964.
Signal Process. (ICASSP), May 2020, pp. 6184–6188. [48] C. Veaux, J. Yamagishi, and K. MacDonald, ‘‘CSTR VCTK corpus:
[27] M. Chen, X. Tan, Y. Ren, J. Xu, H. Sun, S. Zhao, and T. Qin, ‘‘MultiSpeech: English multi-speaker corpus for CSTR voice cloning toolkit,’’ Centre
Multi-speaker text to speech with transformer,’’ in Proc. Interspeech, Speech Technol. Res. (CSTR), Univ. Edinburgh, Tech. Rep., 2019, doi:
Oct. 2020, pp. 4024–4028. 10.7488/ds/2645.
[28] R. Valle, J. Li, R. Prenger, and B. Catanzaro, ‘‘Mellotron: Multispeaker [49] H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and
expressive voice synthesis by conditioning on rhythm, pitch and global Y. Wu, ‘‘LibriTTS: A corpus derived from LibriSpeech for text-to-speech,’’
style tokens,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. 2019, arXiv:1904.02882.
(ICASSP), May 2020, pp. 6189–6193. [50] The LJ Speech Dataset. Accessed: Mar. 15, 2021. [Online]. Available:
[29] A. Kulkarni, V. Colotte, and D. Jouvet, ‘‘Improving latent representation [Link]
for end-to-end multispeaker expressive text to speech system,’’ Tech. [51] CorentinJ/Real-Time-Voice-Cloning: Clone a Voice in 5 Seconds to Gen-
Rep. ffhal-02978485v1f, 2020. erate Arbitrary Speech in Real-Time. Accessed: May 27, 2021. [Online].
[30] E. Cooper, C.-I. Lai, Y. Yasuda, and J. Yamagishi, ‘‘Can speaker aug- Available: [Link]
mentation improve multi-speaker end-to-end TTS?’’ in Proc. Interspeech, [52] L. Mcinnes and J. Healy, ‘‘UMAP: Uniform manifold approximation and
Oct. 2020, pp. 1–5. projection for dimension reduction,’’ 2018, arXiv:1802.03426.
[31] M. Chen, M. Chen, S. Liang, J. Ma, L. Chen, S. Wang, and J. Xiao, ‘‘Cross- [53] Fatchord/WaveRNN: WaveRNN Vocoder + TTS. Accessed: May 27, 2021.
lingual, multi-speaker text-to-speech synthesis using neural speaker [Online]. Available: [Link]
embedding,’’ in Proc. Interspeech, Sep. 2019, pp. 2105–2109, doi: [54] P.-C. Hsu, C.-H. Wang, A. T. Liu, and H.-Y. Lee, ‘‘Towards robust neural
10.21437/Interspeech.2019-1632. vocoding for speech generation: A survey,’’ 2019, arXiv:1912.02461.
[32] W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, [55] J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz,
J. Raiman, and J. Miller, ‘‘Deep voice 3: Scaling text-to-speech with convo- R. Barra-Chicote, A. Moinet, and V. Aggarwal, ‘‘Towards achieving robust
lutional sequence learning,’’ Univ. California, Berkeley, Tech. Rep., 2017. universal neural vocoding,’’ 2018, arXiv:1811.06292.
[33] Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, [56] P. L. Tobing and T. Toda, ‘‘High-fidelity and low-latency universal neural
I. L. Moreno, and Y. Wu, ‘‘Transfer learning from speaker verification to vocoder based on multiband WaveRNN with data-driven linear prediction
multispeaker text-to-speech synthesis,’’ in Proc. Adv. Neural Inf. Process. for discrete waveform modeling,’’ in Proc. Interspeech, 2021, p. 2105.
Syst., Dec. 2018, pp. 4480–4490. [57] D. Paul, Y. Pantazis, and Y. Stylianou, ‘‘Speaker conditional WaveRNN:
[34] R. J. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, Towards universal neural vocoder for unseen speaker and recording condi-
J. Shor, R. Weiss, R. Clark, and R. A. Saurous, ‘‘Towards end-to- tions,’’ in Proc. Interspeech, 2020.
end prosody transfer for expressive speech synthesis with tacotron,’’ in [58] Z. Cai, C. Zhang, and M. Li, ‘‘From speaker verification to multispeaker
Proc. 35th Int. Conf. Mach. Learn., vol. 80. Stockholm, Sweden, 2018, speech synthesis, deep transfer with feedback constraint,’’ in Proc. Inter-
pp. 4693–4702. speech, 2020, pp. 3974–3978, doi: 10.21437/Interspeech.2020-1032.
[59] B. Naderi and R. Cutler, ‘‘An open source implementation of ITU- DAN BIGIOI (Graduate Student Member, IEEE)
T recommendation P.808 with validation,’’ in Proc. Interspeech, 2020, received the bachelor’s degree in electronic and
pp. 2862–2866, doi: 10.21437/Interspeech.2020-2665. computer engineering from the National Univer-
[60] E. H. Rothauser, ‘‘IEEE recommended practice for speech quality mea- sity of Ireland Galway, in 2020. Upon graduat-
surements,’’ IEEE Trans. Audio Electroacoustics, vol. AU-17, no. 3, ing, he worked as a Research Assistant at NUIG
pp. 225–246, Jun. 1969. studying the text-to-speech and speaker recog-
[61] M. Viswanathan and M. Viswanathan, ‘‘Measuring speech quality for text-
to-speech systems: Development and assessment of a modified mean opin- nition methods under the DAVID (Data-Center
ion score (MOS) scale,’’ Comput. Speech Lang., vol. 19, no. 1, pp. 55–83, Audio/Visual Intelligence on-Device) Project.
Jan. 2005, doi: 10.1016/[Link].2003.12.001. Currently, he is working on his Ph.D. at NUIG,
[62] M. Chinen, F. S. C. Lim, J. Skoglund, N. Gureev, and A. Hines, ‘‘ViSQOL sponsored by D-REAL and the SFI Centre for
v3: An open source production ready objective speech and audio metric,’’ Research Training in Digitally Enhanced Reality. His research interests
Tech. Rep., Oct. 2021. include novel deep learning-based techniques for automatic speech dubbing
[63] C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, ‘‘A short- and discovering new ways to process multimodal audio/visual data.
time objective intelligibility measure for time-frequency weighted noisy
speech,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal Process.,
Mar. 2010, pp. 4214–4217, doi: 10.1109/ICASSP.2010.5495701.
[64] T. H. Falk, C. Zheng, and W.-Y. Chan, ‘‘A non-intrusive quality and
intelligibility measure of reverberant and dereverberated speech,’’ IEEE
Trans. Audio, Speech, Language Process., vol. 18, no. 7, pp. 1766–1774,
Sep. 2010, doi: 10.1109/TASL.2010.2052247.
[65] A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, ‘‘Percep-
tual evaluation of speech quality (PESQ)—A new method for speech
quality assessment of telephone networks and codecs,’’ in Proc. IEEE
Int. Conf. Acoust., Speech, Signal Process., May 2001, pp. 749–752, doi:
10.1109/ICASSP.2001.941023. PETER CORCORAN (Fellow, IEEE) is currently
[66] C.-C. Lo, S. W. Fu, W. C. Huang, X. Wang, J. Yamagishi, Y. Tsao, holding the Personal Chair of Electronic Engineer-
and H. M. Wang, ‘‘MOSNet: Deep learning-based objective assessment ing with the College of Science and Engineering,
for voice conversion,’’ in Proc. Interspeech, 2019, pp. 1541–1545, doi: National University of Ireland Galway (NUIG).
10.21437/Interspeech.2019-2003. He was the Co-Founder of several start-up com-
[67] M. Schoeffler, S. Bartoschek, F.-R. Stöter, M. Roess, S. Westphal, B. Edler,
panies, notably FotoNation (currently the Imaging
and J. Herre, ‘‘WebMUSHRA—A comprehensive framework for web-
based listening tests,’’ J. Open Res. Softw., vol. 6, no. 1, p. 8, Feb. 2018, Division, Xperi Corporation). He has more than
doi: 10.5334/jors.187. 600 cited technical publications and patents, more
[68] F. Ribeiro, D. Florencio, C. Zhang, and M. Seltzer, ‘‘CROWDMOS: than 120 peer-reviewed journal articles, 160 inter-
An approach for crowdsourcing mean opinion score studies,’’ in Proc. national conference papers, and a co-inventor on
IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), May 2011, more than 300 granted U.S. patents. He is an IEEE Fellow recognized for
pp. 2416–2419. his contributions to digital camera technologies, notably in-camera red-eye
[69] S. Zielinski, P. Hardisty, C. Hummersone, and F. Rumsey, ‘‘Potential biases correction and facial detection. He is also a member of the IEEE Consumer
in MUSHRA listening tests,’’ in Proc. Audio Eng. Soc. 123rd Audio Eng. Technology Society for more than 25 years and the Founding Editor of IEEE
Soc. Conv., vol. 2, 2007, pp. 1–10. Consumer Electronics Magazine.
[70] M. Wester, C. Valentini-Botinhao, and G. E. Henter, ‘‘Are we using
enough listeners? No! An empirically-supported critique of Interspeech
2014 TTS evaluations,’’ in Proc. Interspeech, 2015, pp, 3476–3480, doi:
10.21437/Interspeech.2015-689.