0% found this document useful (0 votes)
4 views5 pages

DiffCSS: Diverse Conversational Speech Synthesis

The document presents DiffCSS, a novel framework for conversational speech synthesis that utilizes diffusion models and a language model-based TTS backbone to enhance the diversity and expressiveness of synthesized speech. It addresses limitations of existing systems by generating contextually coherent speech with varying prosody, thereby improving naturalness and quality. Experimental results show that DiffCSS significantly outperforms traditional CSS systems in terms of expressiveness and contextual coherence.

Uploaded by

lmsbuddy320
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views5 pages

DiffCSS: Diverse Conversational Speech Synthesis

The document presents DiffCSS, a novel framework for conversational speech synthesis that utilizes diffusion models and a language model-based TTS backbone to enhance the diversity and expressiveness of synthesized speech. It addresses limitations of existing systems by generating contextually coherent speech with varying prosody, thereby improving naturalness and quality. Experimental results show that DiffCSS significantly outperforms traditional CSS systems in terms of expressiveness and contextual coherence.

Uploaded by

lmsbuddy320
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

DiffCSS: Diverse and Expressive Conversational

Speech Synthesis with Diffusion Models


Weihao Wu1,3,∗ , Zhiwei Lin1,∗ , Yixuan Zhou1 , Jingbei Li4 , Rui Niu1,3 , Qinghua Wu3 ,
Songjun Cao3 , Long Ma3 , Zhiyong Wu1,2,†
1
Shenzhen International Graduate School, Tsinghua University, Shenzhen, China
2
The Chinese University of Hong Kong, Hong Kong SAR, China
3
Tencent Youtu Lab 4 StepFun
{wuwh23, lzw22}@[Link], zywu@[Link]
arXiv:2502.19924v1 [[Link]] 27 Feb 2025

Abstract—Conversational speech synthesis (CSS) aims to syn- of more stable and discriminative context representation. Most
thesize both contextually appropriate and expressive speech, and recently, Liu et al. [10] employ a GPT-based architecture to
considerable efforts have been made to enhance the understand- capture semantic and stylistic features within conversational
ing of conversational context. However, existing CSS systems
are limited to deterministic prediction, overlooking the diversity sequences.
of potential responses. Moreover, they rarely employ language However, current CSS methods are constrained by deter-
model (LM)-based TTS backbones, limiting the naturalness and ministic prosody predictions. Specifically, in a given conver-
quality of synthesized speech. To address these issues, in this sational context, speech can be delivered in various ways, each
paper, we propose DiffCSS, an innovative CSS framework that conveying different emotions and intentions. Consequently,
leverages diffusion models and an LM-based TTS backbone to
generate diverse, expressive, and contextually coherent speech. A multiple prosody variations can correspond to the same conver-
diffusion-based context-aware prosody predictor is proposed to sational context. This one-to-many mapping problem has not
sample diverse prosody embeddings conditioned on multimodal been thoroughly explored in previous CSS systems, thus lim-
conversational context. Then a prosody-controllable LM-based iting their ability to generate diverse and expressive prosody.
TTS backbone is developed to synthesize high-quality speech with Additionally, nearly all previous CSS approaches have relied
sampled prosody embeddings. Experimental results demonstrate
that the synthesized speech from DiffCSS is more diverse, on traditional TTS backbones, resulting in limited naturalness
contextually coherent, and expressive than existing CSS systems1 . and quality. Recent advancements in text-modal language
Index Terms—Conversational speech synthesis, diffusion mod- models [11]–[13] have further driven the development of
els, prosody diversity language model (LM)-based TTS systems [14]–[17]. These
systems employ pre-trained audio codec models [18]–[20] to
I. I NTRODUCTION encode speech waveforms into discrete codes, which are then
With the development of deep learning, end-to-end text- predicted by prompt-based language models. When trained
to-speech (TTS) systems have made significant strides [1]– on large-scale datasets, these models can effectively extract
[4]. However, in certain application scenarios such as chatbots semantic information and synthesize speech that closely ap-
and virtual assistants, these systems often underperform due proximates human speech in terms of expressiveness and voice
to their limited ability to understand conversational context, quality.
highlighting the growing importance of conversational speech In this paper, we propose a novel CSS framework: DiffCSS.
synthesis (CSS). CSS aims to generate speech that is not only Inspired by the success of diffusion models in capturing
appropriate for the current utterance but also coherent with complex data distributions and generating diverse outputs in
the broader conversational context. Therefore, effective context the field of computer vision [21]–[24], DiffCSS leverages
modeling is crucial for CSS models. Guo et al. [5] are the diffusion models to generate diverse and expressive prosody
first to propose a GRU-based conversational context encoder conditioned on the conversational context. Additionally, Dif-
that sequentially processes the textual conversational context. fCSS integrates an LM-based TTS backbone to synthesize
Subsequently, graph-based networks [6], [7] are introduced high-quality speech. In this framework, we first employ a pre-
to model the cross-sentence and cross-speaker relationships. trained codec [25] model to extract prosody-related features
To enhance the extraction of fine-grained details, multi-scale from reference speech. Based on these features, we develop
information has also been explored [7], [8]. Deng et al. [9] a prosody-enhanced ParlerTTS [17], which is capable of syn-
introduce contrastive learning to CSS, facilitating the learning thesizing expressive speech conditioned on provided prosody
embeddings. To generate diverse and context-appropriate
This work is supported by National Natural Science Foundation prosody embeddings from multimodal conversational context,
of China (62076144) and Shenzhen Science and Technology Program we design a diffusion-based context-aware prosody predictor.
(WDZC20220816140515001, JCYJ20220818101014030). Experimental results demonstrate that our proposed method
*These authors contributed equally to this work as first authors.
† Corresponding Author. significantly outperforms deterministic baselines in terms of
1 Audio samples: [Link] expressiveness and contextual coherence. Furthermore, the
Decoder-only Transformer
0 0 0 t1 t2 … tn-3

0 0 t1 t2 t3 … tn-3 tn-2
Current Text RVQ
3059 6 24 …
“Tom, that lucky guy.” 0 t1 t2 t3 t4 … tn-2 tn-1 Decoder
Pre-pend tn-1
Synthesized Speech
t1 t2 t3 t4 t5 … tn
text tokens
Cross Attention
Speaker Embedding
+

Denoising Diffusion-based
Prosody Embedding P1 P2 … Pm context-aware …
prosody predictor
Gauss Noise
Cross Attention
K,V
FACodec F1 F2 … Fn QKV-Attn

Reference Speech Prosody Features Q


Text M: I have some good news for you.
Q1 Q2 … Qm W: What’s that?
concat M: Jenny is getting Married.
Prosody extractor Learnable Queries W: Great! Who is the bridegroom?
Speech
Prosody-enhanced ParlerTTS Multi-modal Conversational Context

Fig. 1. Overall architecture of the proposed CSS framework

prosody distribution generated by DiffCSS aligns more closely reference speech. To extract prosody features disentangled
with ground truth distribution, highlighting the effectiveness of from other information, we employ the pre-trained FACodec
incorporating diffusion models into CSS. [25] to extract frame-level prosody features {F1 , F2 , ..., Fn },
To summarize, the main contributions of this paper are: where n represents the number of audio frames. However,
• We proposed a novel CSS framework DiffCSS, which the computation cost of the TTS backbone increases linearly
leverages diffusion models to enhance prosody diversity with the length of prosody features. To reduce resource con-
depending on conversational context along with an LM- sumption, we introduce a cross-attention layer with learnable
based TTS backbone to improve speech quality. query tokens {Q1 , Q2 , ..., Qm } to derive fixed-length prosody
• To the best of our knowledge, we are the first to apply dif- embeddings {P1 , P2 , ..., Pm }, where m ≪ n.
fusion models to CSS systems for conversational context These prosody embeddings are then combined with pre-
modeling. Through prosody diffusion, we significantly extracted speaker embeddings and serve as the keys and values
enhance the prosody diversity of the synthesized speech. in the cross-attention layer of the TTS backbone, guiding the
• Through comparative experiments, we demonstrate the speech synthesis process.
effectiveness of our proposed approach in generating
speech that is both rich in prosody diversity and suitable B. Diffusion-based context-aware prosody predictor
for the conversational context. To predict diverse and contextually appropriate prosody, we
II. M ETHODOLOGY design a diffusion-based context-aware prosody predictor that
generates the current prosody embedding conditioned on both
The architecture of our proposed model is illustrated in Fig. the multimodal conversational context and the current text.
1. It consists of two primary components: a TTS backbone The structure of the prosody predictor is illustrated in Fig. 2,
based on ParlerTTS [17] and a diffusion-based context-aware where a set of Transformer encoder blocks are employed as
prosody predictor. The TTS backbone synthesizes high-quality the underlying denoiser network θ.
speech based on varying prosody inputs, while the prosody For a conversation chunk of length N + 1, we combine the
predictor generates diverse prosody embeddings conditioned multimodal context information as described in Equation 1,
on the conversational context. where c denotes the multimodal context information, si and
A. Prosody-enhanced ParlerTTS pi denote the textual information extracted by a pre-trained
sentence-level T5 [27] and the prosody embedding for the i-
Inspired by the success of LM-based TTS models, we th turn, respectively.
develop a prosody-enhanced ParlerTTS [17] as our TTS back-
bone, which predicts pre-extracted acoustic tokens by decoder- c = [s1 , p1 , ..., sN , pN ] (1)
only transformer blocks. By adopting the delayed pattern
introduced in [26], the TTS backbone is capable of generating During the diffusion process, Gaussian noise is added to the
high-quality speech autoregressively in a short amount of time. current prosody embedding pN +1 according to a fixed noise
Within the backbone, we designed a prosody extractor to schedule α1 , ..., αT , where T is the total diffusion steps. This
learn prosody embeddings in an unsupervised manner from process can be described as Equation 2 and 3, where ϵ ∼
× 𝑁𝑁 𝑏𝑏𝑏𝑏𝑏𝑏𝑏𝑏𝑏𝑏𝑏𝑏 During inference, we first extract the multimodal conversa-
zT z0
tional context information c and the current textual information

Cross Attention

Feed Forward
Self Attention
… …
sN +1 . We then iteratively compute the predicted prosody
Gauss Noise Prosody embedding embedding zˆ0 from Gaussian noise, following Equation 5 and
6. Finally, we feed zˆ0 into the TTS backbone and synthesize
denoising × 𝑇𝑇 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠

ST5
speech according to it.
Concat III. E XPERIMENTS
Text A. Training setup
N+1th turn
Prosody Prosody We conduct our experiments on two English datasets. For
ST5 ST5 embedding
embedding
… TTS backbone pre-training, we use an open-source multi-
Text Speech Text Speech speaker English speech dataset LibriTTS-R [28], which con-
1st turn Nth turn tains 585 hours of high-quality speech from 2,456 speakers.
Fig. 2. The structure of the diffusion-based context-aware prosody predictor For the conversational dataset, we select DailyTalk [29],
which includes 2,541 conversations spanning about 20 hours,
N (0, I) is the Gaussian noise, z0 represents the ground truth performed by two speakers. These conversations are divided
n
prosody embedding pN +1 and αt =
Q
αi . into equal-length chunks, each containing 5 utterances, with
i=1 the first 4 serving as the conversational context for the fifth.
√ The first 2400 conversations are used for training while the
q(zt | zt−1 ) = N (zt ; αt zt−1 , (1 − αt )I) (2)
√ √ remaining 141 are reserved for testing. A pre-trained audio
zt = αt z0 + 1 − αt ϵ (3) codec model DAC [20] is employed to encode the raw wave-
During training, we first uniformly sample a time step t from form with 24kHz sampling rate and reconstruct the waveform
[1, T], based on which the ground truth prosody embedding based on the predicted acoustic tokens. Speaker embeddings
z0 is diffused into zt according to Formula 3. Then we are extracted by a pre-trained voiceprint model2 .
concatenate textual information of the current sentence sN +1 In our implementation, the TTS backbone comprises 12
with zt as the input of the diffusion denoiser. Multimodal transformer decoder blocks, and the diffusion-based context-
context information c is incorporated through cross-attention. aware prosody predictor consists of 6 transformer encoder
Given zt , sN +1 , t and c, the denoiser predicts the added blocks. We pre-train the TTS backbone on LibriTTS-R for
Gaussian noise ϵθ (zt , sN +1 , t, c) and optimizes its parameters 150000 iterations and finetune it on DailyTalk for 20000
with the following denoising objective: iterations with a batch size of 64. The diffusion-based context-
aware prosody predictor is trained on DailyTalk for 100000
L(θ) = Et,zt ,ϵ ∥ϵ − ϵθ (zt , sN +1 , t, c)∥2 (4) iterations with a batch size of 32.
In the inference stage, we sample zT ∼ N (0, I) and itera- B. Baseline Models
tively perform the denoising process according to Equation 5 We implemented the following three models as baselines.
and 6 to compute zt−1 from zt for t = T, T − 1, ..., 1 where GRU-based context modeling [5] This model uses a unidi-
x ∼ N (0, I) except for x = 0 when t = 1. After completing rectional GRU to process conversational context sequentially.
T iterations, we obtain the generated prosody embedding zˆ0 . DialogueGCN-based context modeling [6] This model
1

1 − αt
 employs a relation-aware graph convolutional network to
zt−1 = √ xt − √ ϵθ (zt , sN +1 , t, c) + σt x (5) model context from both text and audio modalities.
at 1 − αt
Transformer encoder-based context modeling We imple-
1 − αt−1
σt = (1 − αt ) (6) ment a contextual prosody predictor based on the Transformer
1 − αt
encoder, which shares the same structure and parameters as our
C. Training Strategy and Inference Procedure proposed method but without diffusion modeling.
To enhance the controllability of the TTS module and C. Subjective Evaluation
improve the prosody diversity of synthesized speech, we pro-
pose a two-stage training strategy, where prosody is explicitly To evaluate the performance of our proposed model in com-
modeled as an intermediate representation. In order to enhance parison to baselines, we conduct two separate mean opinion
the semantic understanding and speech synthesis capability of score (MOS) tests: one for speech expressiveness and the other
the TTS backbone, we first pre-train the TTS backbone on a for contextual coherence. We randomly select 15 samples and
large-scale dataset and subsequently fine-tune it on a conver- their corresponding contexts from the test set for evaluation. A
sational dataset. After completing the TTS backbone training, total of 20 listeners are invited to rate the synthesized speech
we freeze its parameters and use it to extract ground truth on a scale of 1 to 5 with 1 point interval, based on both
prosody embeddings from reference speech. We then divide expressiveness and contextual coherence.
the conversations into equal-length chunks, which are used to 2 Available at: [Link]
train the diffusion-based context-aware prosody predictor. 3dspeaker/sv-cam++
TABLE I
T HE OBJECTIVE AND SUBJECTIVE COMPARISONS FOR DIFFERENT MODELS .

Context modeling method MOS(Expressiveness) ↑ MOS(Coherence) ↑ MCD ↓ NDB ↓ JSD ↓


GRU-based [5] 3.209 ± 0.108 3.177 ± 0.087 8.011 16 0.227
DialogueGCN-based [6] 3.347 ± 0.104 3.362 ± 0.093 7.867 13 0.156
Transformer Encoder-based 3.264 ± 0.112 3.253 ± 0.103 7.892 14 0.181
Proposed 3.602 ± 0.101 3.574 ± 0.096 7.745 4 0.036

As shown in Table I, our proposed method achieves the best


E-MOS of 3.602 and C-MOS of 3.574 compared to baselines.
This indicates that our model can generate prosody that is
both contextually appropriate and expressive. While the Trans-
former encoder-based method outperforms the GRU-based
method, it lags behind the DialogueGCN-based method and
significantly underperforms when compared to the proposed
method. This suggests that although the Transformer encoder
exhibits a moderate ability in modeling conversational context,
the integration of diffusion models is essential for improving
both prosody prediction and contextual comprehension.
D. Objective Evaluation Fig. 3. Distribution of prosody embeddings in predictions and ground truth.
For objective evaluation, we employ Mel-Cepstral Distor- Each unique color corresponds to a specific cluster.
tion (MCD) to assess the overall quality of the synthesized E. Ablation Study on multimodal context
speech, and Number of Statistically-Different Bins (NDB)
along with Jensen-Shannon Divergence (JSD) [30] to evaluate Furthermore, we investigate the impact of multimodal con-
the diversity of prosody, following [31]. NDB and JSD are text by evaluating three different settings: (1) Proposed model
computed through a clustering process. The procedure is as w/o textual context (2) Proposed model w/o acoustic con-
follows: 1) Cluster the ground-truth prosody into n bins to text (3) Proposed model without full context. The modalities
obtain the ground-truth prosody distribution across bins. 2) are excluded by setting the corresponding contextual informa-
Generate prosody samples using the prosody predictor, and tion to zeros. We conducted a separate training session for
assign each generated sample to the closest bin. 3) Calculate each ablation setting, effectively preventing the model from
the proportion of generated samples in each bin, resulting in receiving unwanted contextual information.
a new distribution across the bins. 4) Evaluate the similarity As shown in Table II, the absence of either modality reduces
between the generated and ground-truth distributions. JSD is the overall quality and prosody diversity of the synthesized
the Jensen-Shannon divergence between the two distributions, speech, thereby diminishing the effectiveness of context mod-
and NDB counts the number of bins with statistically signifi- eling. Notably, the exclusion of the acoustic context leads
cant differences in sample proportions. In this evaluation, we to a greater decline in performance compared to the textual
set the number of prosody clustering bins to 20. context, underscoring the crucial role of acoustic information
As presented in Table I, our method achieves an NDB of 4 in modeling conversational context.
and a JSD of 0.036, significantly outperforming all baselines. TABLE II
A BLATION STUDIES ON CONTEXT MODALITIES
This demonstrates that the generated prosody from our method
aligns much more closely with the ground truth prosody MCD ↓ NDB ↓ JSD ↓
distribution. Furthermore, our proposed method also surpasses Proposed 7.745 4 0.036
all baselines in MCD, indicating its ability to synthesize high- Proposed w/o textual context 7.791 5 0.041
Proposed w/o acoustic context 7.946 10 0.129
quality speech. Proposed w/o full context 8.038 12 0.152
We further visualized the prosody distributions of different
methods, as shown in Figure 3, where each unique color IV. C ONCLUSION
corresponds to a specific cluster. The Transformer encoder-
In this paper, we introduce DiffCSS, a novel CSS framework
based method predicts prosody embeddings concentrated in
designed to generate diverse and high-quality speech. We
only a few clusters, with sparse representation in others.
propose a diffusion-based context-aware prosody predictor
In contrast, prosody embeddings generated by our proposed
to generate diverse prosody embeddings conditioned on the
method are more evenly distributed across clusters, closely
conversational context. Additionally, we develop an LM-based
resembling the ground truth distribution. This indicates that
TTS backbone to synthesize high-quality speech based on
the deterministic baseline tends to predict similar, less diverse
sampled prosody embeddings. Experimental results demon-
prosody, while our proposed method exhibits significantly
strate that our proposed model outperforms existing baselines
improved prosody diversity. These results further highlight
in terms of expressiveness, contextual coherence, and prosody
the importance of incorporating diffusion models in enhancing
diversity.
prosody diversity.
R EFERENCES [17] Dan Lyth and Simon King, “Natural language guidance of high-
fidelity text-to-speech with synthetic annotations,” arXiv preprint
[1] Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J arXiv:2402.01912, 2024.
Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, [18] Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and
Samy Bengio, et al., “Tacotron: Towards end-to-end speech synthesis,” Marco Tagliasacchi, “Soundstream: An end-to-end neural audio codec,”
in Proceedings of the Annual Conference of the International Speech IEEE/ACM Transactions on Audio, Speech, and Language Processing,
Communication Association, INTERSPEECH, 2017, pp. 7471 – 7480. vol. 30, pp. 495–507, 2021.
[2] Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep [19] Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi
Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Adi, “High fidelity neural audio compression,” arXiv preprint
Rj Skerrv-Ryan, et al., “Natural tts synthesis by conditioning wavenet on arXiv:2210.13438, 2022.
mel spectrogram predictions,” in ICASSP 2018-2018 IEEE International [20] Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar,
Conference on Acoustics, Speech and Signal Processing (ICASSP). and Kundan Kumar, “High-fidelity audio compression with improved
IEEE, 2018, pp. 4779–4783. rvqgan,” in Advances in Neural Information Processing Systems, 2024,
[3] Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie- vol. 36.
Yan Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” [21] Jonathan Ho, Ajay Jain, and Pieter Abbeel, “Denoising diffusion
in International Conference on Learning Representations, 2021. probabilistic models,” in Advances in neural information processing
[4] Jaehyeon Kim, Jungil Kong, and Juhee Son, “Conditional variational systems, 2020, vol. 33, pp. 6840–6851.
autoencoder with adversarial learning for end-to-end text-to-speech,” in [22] Prafulla Dhariwal and Alexander Nichol, “Diffusion models beat gans on
Proceedings of the 38th International Conference on Machine Learning. image synthesis,” in Advances in neural information processing systems,
2021, vol. 139, pp. 5530–5540, PMLR. 2021, vol. 34, pp. 8780–8794.
[5] Haohan Guo, Shaofei Zhang, Frank K Soong, Lei He, and Lei Xie, [23] Alexander Quinn Nichol and Prafulla Dhariwal, “Improved denoising
“Conversational end-to-end tts for voice agents,” in 2021 IEEE Spoken diffusion probabilistic models,” in Proceedings of the 38th International
Language Technology Workshop (SLT). IEEE, 2021, pp. 403–409. Conference on Machine Learning. PMLR, 2021, vol. 139, pp. 8162–
[6] Jingbei Li, Yi Meng, Chenyi Li, Zhiyong Wu, Helen Meng, Chao Weng, 8171.
and Dan Su, “Enhancing speaking styles in conversational text-to-speech [24] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser,
synthesis with graph-based multi-modal context modeling,” in ICASSP and Björn Ommer, “High-resolution image synthesis with latent diffu-
2022-2022 IEEE International Conference on Acoustics, Speech and sion models,” in Proceedings of the IEEE/CVF conference on computer
Signal Processing (ICASSP). IEEE, 2022, pp. 7917–7921. vision and pattern recognition, 2022, pp. 10684–10695.
[7] Jingbei Li, Yi Meng, Xixin Wu, Zhiyong Wu, Jia Jia, Helen Meng, [25] Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao
Qiao Tian, Yuping Wang, and Yuxuan Wang, “Inferring speaking styles Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, et al.,
from multi-modal conversational context by multi-scale relational graph “NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and
convolutional networks,” in Proceedings of the 30th ACM International diffusion models,” in Proceedings of the 41st International Conference
Conference on Multimedia, 2022, pp. 5811–5820. on Machine Learning. 2024, vol. 235, pp. 22605–22623, PMLR.
[8] Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li, Yingming Gao, [26] Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel
Jianhua Tao, Jianqing Sun, and Jiaen Liang, “M 2-ctts: End-to-end multi- Synnaeve, Yossi Adi, and Alexandre Défossez, “Simple and controllable
scale multi-modal conversational text-to-speech synthesis,” in ICASSP music generation,” in Advances in Neural Information Processing
2023-2023 IEEE International Conference on Acoustics, Speech and Systems, 2024, vol. 36.
Signal Processing (ICASSP). IEEE, 2023, pp. 1–5. [27] Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith B
[9] Yayue Deng, Jinlong Xue, Yukang Jia, Qifei Li, Yichen Han, Fengping Hall, Daniel Cer, and Yinfei Yang, “Sentence-t5: Scalable sentence
Wang, Yingming Gao, Dengfeng Ke, and Ya Li, “Concss: Contrastive- encoders from pre-trained text-to-text models,” in Findings of the
based context comprehension for dialogue-appropriate prosody in con- Association for Computational Linguistics: ACL 2022, 2022, pp. 1864–
versational speech synthesis,” in ICASSP 2024-2024 IEEE International 1874.
Conference on Acoustics, Speech and Signal Processing (ICASSP). [28] Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding, Kohei Yatabe,
IEEE, 2024, pp. 10706–10710. Nobuyuki Morioka, Michiel Bacchiani, Yu Zhang, Wei Han, and Ankur
[10] Rui Liu, Yifan Hu, Ren Yi, Yin Xiang, and Haizhou Li, “Generative Bapna, “Libritts-r: A restored multi-speaker text-to-speech corpus,”
expressive conversational speech synthesis,” in Proceedings of the 32nd in Proceedings of the Annual Conference of the International Speech
ACM International Conference on Multimedia, 2024, pp. 4187 – 4196. Communication Association, INTERSPEECH, 2023, pp. 5496 – 5500.
[11] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, [29] Keon Lee, Kyumin Park, and Daeyoung Kim, “Dailytalk: Spoken
“Bert: Pre-training of deep bidirectional transformers for language dialogue dataset for conversational text-to-speech,” in ICASSP 2023-
understanding,” in Proceedings of the 2019 Conference of the North 2023 IEEE International Conference on Acoustics, Speech and Signal
American Chapter of the Association for Computational Linguistics: Processing (ICASSP). IEEE, 2023, pp. 1–5.
Human Language Technologies, Volume 1 (Long and Short Papers), [30] Eitan Richardson and Yair Weiss, “On gans and gmms,” in Advances
2019, pp. 4171–4186. in neural information processing systems, 2018, vol. 31.
[12] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared [31] Rongjie Huang, Chunlei Zhang, Yi Ren, Zhou Zhao, and Dong Yu,
Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish “Prosody-tts: Improving prosody with masked autoencoder and condi-
Sastry, Amanda Askell, et al., “Language models are few-shot learners,” tional diffusion model for expressive text-to-speech,” in Findings of the
in Advances in Neural Information Processing Systems, 2020. Association for Computational Linguistics: ACL 2023, 2023, pp. 8018–
[13] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan 8034.
Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu, “Explor-
ing the limits of transfer learning with a unified text-to-text transformer,”
Journal of machine learning research, vol. 21, no. 140, pp. 1–67, 2020.
[14] Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou,
Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al.,
“Neural codec language models are zero-shot text to speech synthesiz-
ers,” arXiv preprint arXiv:2301.02111, 2023.
[15] Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu,
Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al.,
“Speak foreign languages with your own voice: Cross-lingual neural
codec language modeling,” arXiv preprint arXiv:2303.03926, 2023.
[16] Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel
Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar,
et al., “Voicebox: Text-guided multilingual universal speech generation
at scale,” in Advances in neural information processing systems, 2024,
vol. 36.

You might also like