0% found this document useful (0 votes)
27 views5 pages

Single-Codec: Efficient Speech Codec

The document presents Single-Codec, a novel single-codebook speech codec designed to enhance efficiency and performance in text-to-speech (TTS) applications. It utilizes a disentangled VQ-VAE architecture with several key components, including a global reference encoder and a hybrid sampling module, to improve speech generation quality while maintaining a low bandwidth of 304bps. Experimental results demonstrate that Single-Codec outperforms existing multi-codebook codecs in terms of reconstruction quality, naturalness, and intelligibility in TTS systems.

Uploaded by

hanklee508
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
27 views5 pages

Single-Codec: Efficient Speech Codec

The document presents Single-Codec, a novel single-codebook speech codec designed to enhance efficiency and performance in text-to-speech (TTS) applications. It utilizes a disentangled VQ-VAE architecture with several key components, including a global reference encoder and a hybrid sampling module, to improve speech generation quality while maintaining a low bandwidth of 304bps. Experimental results demonstrate that Single-Codec outperforms existing multi-codebook codecs in terms of reconstruction quality, naturalness, and intelligibility in TTS systems.

Uploaded by

hanklee508
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Single-Codec: Single-Codebook Speech Codec towards High-Performance

Speech Generation

Hanzhao Li1 , Liumeng Xue2 , Haohan Guo3 , Xinfa Zhu1 , Yuanjun Lv1 , Lei Xie1,∗ ,
Yunlin Chen4 , Hao Yin4 , Zhifei Li4
1
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science,
Northwestern Polytechnical University, Xi’an, China
2
School of Data Science, The Chinese University of Hong Kong,
Shenzhen (CUHK-Shenzhen), China
3
The Chinese University of Hong Kong, Hong Kong SAR, China
4
Shanghai Mobvoi Information Technology Co., Ltd
lihanzhao@[Link], lxie@[Link]
arXiv:2406.07422v1 [[Link]] 11 Jun 2024

Abstract embeddings in the LLM of predicted speech units. These em-


The multi-codebook speech codec enables the applica- beddings contain more information related to the input text
tion of large language models (LLM) in TTS but bottlenecks and the target speaker to compensate for the compression loss.
efficiency and robustness due to multi-sequence prediction. However, this operation introduces more training and inference
To avoid this obstacle, we propose Single-Codec, a single- costs. Recently, TiCodec [6] proposes introducing an additional
codebook single-sequence codec, which employs a disentan- global encoder to disentangle time-invariant information out of
gled VQ-VAE to decouple speech into a time-invariant embed- speech units, reducing the amount of frame-level information
ding and a phonetically-rich discrete sequence. Furthermore, that needs encoding. It inspires us to re-think speech codec from
the encoder is enhanced with 1) contextual modeling with a the perspective of feature disentanglement.
BLSTM module to exploit the temporal information, 2) a hybrid In this study, we propose a single-codebook neural audio
sampling module to alleviate distortion from upsampling and codec, Single-Codec, for high-performance speech generation.
downsampling, and 3) a resampling module to encourage dis- Single-Codec performs compression and reconstruction on Mel
crete units to carry more phonetic information. Compared with Spectrogram instead of the raw waveform, enabling efficient
multi-codebook codecs, e.g., EnCodec and TiCodec, Single- compression of speech information while preserving important
Codec demonstrates higher reconstruction quality with a lower details, as stated in Tortoise-TTS [10]. To further enhance
bandwidth of only 304bps. The effectiveness of Single-Code is the codec performance and applicability to speech synthesis,
further validated by LLM-TTS experiments, showing improved Single-Codec incorporates several key components:
naturalness and intelligibility. • A global reference encoder to decouple time-invariant fea-
Index Terms: speech codec, single-codebook codec, language tures. Specifically, we utilize continuous global represen-
model, text-to-speech tations rather than discrete representations and longer refer-
ence segments to capture more acoustic details, enabling em-
1. Introduction bedding sufficient phonetic information into single-codebook
Large language models (LLMs) have attracted wide attention discrete units.
in the speech domain, particularly in text-to-speech synthesis • A BLSTM module for contextual modeling to help discover
(TTS) [1, 2]. In such LLM-based TTS systems, to oper- the correlation between adjacent frames, enhancing speech
ate speech synthesis as a simple next-token prediction prob- content clustering efficiency.
lem, the first thing is to seek an appropriate speech codec • A hybrid sampling module that uses both convolution and
for speech tokenization and waveform reconstruction. Multi- pooling to achieve downsampling, and transposed convolu-
codebook codecs [3], as the SOTA approaches, are widely tion and replication to achieve upsampling, alleviating up-
adopted in LLM-based TTS to achieve superior reconstruction sampling and downsampling distortion.
quality. However, they also require the LLM to predict multiple
• A resampling module to encourage the encoder to extract
discrete sequences, affecting efficiency and stability seriously,
more phonetics-relevant information with lower short-time
although various designs of codec [4, 5, 6] and LLM [7, 8, 9]
variance from the acoustic sequence.
are proposed to adapt this multi-sequence discrete representa-
tion better. Hence, seeking an effective approach to obtain the To the best of our knowledge, Single-Codec is the first
single-sequence discrete speech representation is critical to by- single-codebook codec dedicatedly designed for LLM-based
pass this limitation. speech generation. We comprehensively compare Single-
However, unlike the text, it is impossible to completely rep- Codec with SOTA multi-codebook codecs, including EnCodec
resent the speech audio with abundant information in seman- and TiCodec, by conducting both objective and subjective
tics and acoustics with only one discrete token sequence. Al- tests in analysis-synthesis and TTS. The results show that
though Tortoise-TTS [10] achieves the LLM with the single- Single-Codec with the lower bandwidth has better speech
sequence discrete speech representation, A diffusion model still reconstruction quality, and significantly improves intelligi-
needs to be trained to generate Mel Spectrograms from latent bility, naturalness, and speaker similarity of synthesized
speech in zero-shot LLM-TTS. Audio samples are available at
* Corresponding author. [Link]
2. Method Addition

Subtraction
2.1. Architecture of Single-Codec
FFT Pre-trained
The architecture of Single-Codec is shown in Figure 1. It is built
on Vector Quantised-Variational AutoEncoder (VQVAE) [11]
with Mel Spectrogram input and reconstruction, similar to Tor- Conformer Block
toise TTS [10]. We adopt a Conformer-based encoder to encode Hybird sampling
a Mel Spectrogram segment seg2 into a latent content repre-
Conformer Block
sentation c, which is then passed to the Vector Quantizer (VQ) Encoder
Hybird sampling
for vector quantization. The convolution-based decoder recon- Conformer
Encoder
structs the Mel Spectrogram seg ˜ 2 from the quantized content Comformer Block

representation c. Additionally, we apply a discriminator to im-


Resampling
prove generation quality [12]. Finally, we use a neural vocoder Module

BigVGAN [13] to reconstruct waveform from codec output,


Downsample
i.e., Mel Spectrogram. BLSTM Module
Reference Encoder
Upsample
To achieve a high-quality single-codebook codec, we im- Conv2d
Residual Connect
prove the codec architecture with four modules. Specifically, VQ
GRU
we add a reference encoder to decouple time-invariant informa-
tion in speech from a Mel Spectrogram segment seg1 , yield-
BLSTM Module
ing a global representation g. A hybrid sampling module is Residual Block

adopted to alleviate sampling loss. Moreover, we introduce


Resampling Hybird sampling
a BLSTM [14] module and resampling module in both codec Module
Residual Block
encoder and decoder to enhance contextual information and
Hybird sampling
phonetics-relevant information, respectively. Conv
Decoder Residual Block

Decoder
2.2. Reference Encoder
Real
Speech contains multiple aspects of information, such as time-
Discriminator BigVGAN
variant content, time-invariant timbre, and acoustic environ-
Fake
ment. Multiple codebook in codec makes it easy to encode these
various information. However, for a single-codebook codec, it
is challenging to compress all information into a limited num- Figure 1: The architecture of Single-Codec.
ber of discrete units. To solve this problem, we decouple global
information (such as timbre and acoustic environment) that is
almost invariable in all frames of the utterance and discretize
speech content into code.
2.4. Hybrid Sampling Module
We introduce a reference encoder to derive global repre-
sentation g that is mainly related to timbre and acoustic envi- Neural codecs usually employ a sampling module to reduce the
ronment. The input of the reference encoder is a segment seg1 sequence length of the discrete representation. Currently, the
randomly selected from the input utterance. We set the length up-sampling and down-sampling operations in codecs are usu-
of the segment seg1 for reference input to 600 frames while the ally implemented by convolution, transposed convolution, or
input segment seg2 for codec encoder to 200 frames, where the pooling and repeat. The sampling process inevitably produces
short segment seg2 can reduce the amount of calculation and sampling loss, resulting in reduced encoding and decoding ca-
memory overhead, while the longer reference segment seg1 can pabilities. Inspired by MR-HuBERT [17], we introduce an im-
help to obtain more robust global features. The output g of the proved hybrid sampling module that uses both convolution and
reference encoder is fed to the codec encoder and decoder after pooling to achieve downsampling and transposed convolution
passing through different linear layers, where it subtracts with and replication to achieve upsampling. The combination of dif-
output of the encoder blocks and adds to the input of the decoder ferent sampling methods can alleviate sampling distortion.
blocks.

2.3. BLSTM Module 2.5. Resampling Module

Codecs are generally trained on large-scale speech data to en- The main goal of a single-codebook speech codec is to extract
sure good generalization. The diversity of speech content cre- short-term invariant speech units from acoustic representations.
ates challenges for single-codebook codecs with appropriate The diversity of acoustic representations brings challenges to
sizes. Unlike EnCodec [3], which introduces the sequence the learning of codebook vectors. To solve this problem, we
modelling with LSTM [15] and finds that it can improve the propose a novel resampling module, which first downsamples
Scale-Invariant Signal-to-Noise Ration (SI-SNR) [16], we add the input feature for local modelling and then residual connect
BLSTM modules before and after the quantizer to enhance con- after upsampling. This bottlenecking operation along the time
textual information. We found this can improve the efficiency axis encourages the encoder to extract more phonetics-relevant
of speech content modelling and make it easier to form stable information with lower short-time variance from the acoustic
clustering centers. sequence.
Table 1: Objective metrics scores of various codecs, where the bold numbers highlight best results in 304bps.

Experiment Model Bandwidth (bps) STOI↑ PESQ↑ MCD↓ UTMOS↑ SPK↑


VQVAE 0.805 1.591 4.489 2.847 0.727
Ref-short 0.827 1.797 4.184 2.884 0.790
Ref-long 0.833 1.819 4.122 2.983 0.809
Ablation Ref-BLSTM 304 0.837 1.879 4.054 2.933 0.814
Ref-HybSam 0.837 1.899 4.077 2.961 0.800
Ref-BLSTM-HybSam 0.841 1.912 4.042 3.081 0.809
Ref-BLSTM-HybSam-Conf 0.839 1.901 4.091 3.029 0.811
EnCodec-1VQ 750 0.765 1.319 4.515 2.521 0.599
Comparison TiCodec-1VQ 750 0.808 1.694 4.486 2.919 0.684
TiCodec-2VQ 1500 0.866 2.140 3.865 3.079 0.762
Single-Codec 304 0.842 1.933 4.017 3.031 0.817

3. Experiments block is 1024. The reference encoder consists of 6 layers of 2D


3.1. Dataset convolution with a kernel size of 3 and a GRU [24] layer. The
We train speech codecs and a zero-shot TTS system, VALL- residual block [25] consists of two residual units. Each residual
E [1], using five open-source datasets, including LibriTTS[18], unit includes 2 one-dimensional convolutions with kernel sizes
Hi-Fi TTS[19], VCTK[20], AISHELL-1[21], and AISHELL- of 3 and 1 respectively. Discriminator consists of 4 layers of 2D
3[22]. A total of 1165.3 hours of English and Chinese speech is convolution with a kernel size of 5 and 2 layers of 2D convolu-
used. tion with a kernel size of 3. The BLSTM module contains two
LSTM layers with a hidden size of 128.
3.2. Comparison Models
During training, we conduct 300k iterations using a batch
We adopt EnCodec [3] with one codebook (EnCodec-1VQ) and size of 1024 on a single V100 GPU for Single-Codec. The
TiCodec [6] with one codebook (TiCodec-1VQ) and two code- baseline model Encodec utilizes with the code reimplemented
books (TiCodec-2VQ) as the baselines to compare with our in HifiCodec1 [26], and is trained for 25 epochs. TiCodec2 is
proposed Single-Codec. For VALL-E, we use EnCodec with trained for 300k steps with a batch size of 40 on two V100
one, four, and eight codebooks, representing EnCodec-1VQ, GPUs. For TTS, we employ VALL-E, reimplemented in Am-
EnCodec-4VQ, and EnCodec-8VQ, TiCodec with one code- phion3 [27], with dynamic batch sizing and a maximum token
book (TiCodec-1VQ) as the baselines to evaluate the perfor- limit of 4000 per batch. The single-codebook codec only uti-
mance of codecs on speech synthesis. lizes the AR stage, while the multiple-codebook codec trains
To verify the effectiveness of our designed modules in both the AR and NAR stages simultaneously. Eight A800 GPUs
Single-Codec, we conduct ablation studies on the following and 70 epochs are employed for training VALL-E.
models.
3.4. Ablation Studies
• VQVAE: A basic VQVAE codec with a discriminator for per-
ceptual loss, the structure and configuration of the VQVAE We calculate STOI [28], PESQ [29], Mel cepstral distor-
codec is similar to that in Tortoise TTS [10]. tion (MCD) [30], UTMOS4 [31] and speaker cosine similarity
• Ref-short: VQVAE with a reference encoder that consumes (SPK)5 to objectively evaluate the quality of speech reconstruc-
a short segment with 200 frames as input. tion. The test set is composed of 100 randomly selected sen-
tences from unseen speakers. The objective result is shown in
• Ref-long: VQVAE with a reference encoder that consumes a
Table 1.
long segment with 600 frames as input.
Compared with VQVAE, either Ref-short or Ref-long ob-
• Ref-BLSTM: Ref-long with the BLSTM module to verify
tains better performance on all metrics. It indicates that it
the effectiveness of the BLSTM module.
is effective to decouple global information from speech for
• Ref-HybSam: Ref-long with the hybrid sampling module to the single-codebook codec. Moreover, Ref-long outperforms
verify the effectiveness of the hybrid sampling module. Ref-short in both reconstruction and speaker similarity, sug-
• Ref-BLSTM-HybSam: Ref-long with the BLSTM and hy- gesting that longer reference segments help capture more ac-
brid sampling modules to verify the effectiveness of the com- curate time-invariant information and enhance content mod-
bination of BLSTM and hybrid sampling modules. elling. Ref-BLSTM, Ref-HybSam, and Ref-BLSTM-HybSam
• Ref-BLSTM-HybSam-Conf: Ref-BLSTM-HybSam with get higher reconstruction quality, showing the effectiveness of
the Conformer-based encoder, excluding the resampling the BLSTM and hybrid sampling modules. Moreover, Ref-
module. BLSTM-HybSam-Con yields on-pair performance with Ref-
BLSTM-HybSam but gets further improvement after adding the
3.3. Model Parameters and Training Details resampling module, i.e., our proposed Single-Code, achieving
The audio sample rate is 24khz, and the hop length and window the best results.
length of the Mel Spectrogram are 256 and 1024, respectively. 1 [Link]
The downsample rate is 4, resulting in a total downsampling of 2 [Link]
1024 times (about 23 discrete tokens per second). The code- 3 [Link]
book size is 8192. The model size of the codec is 256. The 4 [Link]
sizes of the intermediate hidden states in convolution blocks are 5 We adopt WeSpeaker [32] to extract speaker embedding for cosine
256, 512, and 1024, while the hidden size of the Conformer [23] similarity compution
0.40
Table 2: The objective and subjective metrics scores of VALL-E.
VQVAE
Ref-short
0.35 Ref-long Subjective Objective
Ref-HybSam Codec Model
Ref-BLSTM N-MOS↑ S-MOS↑ WER↓ S-Cosine↑
0.30
Ref-BLSTM-HybSam VQVAE 2.81 ± 0.06 3.24 ± 0.07 30.9 0.649
Ref-BLSTM-HybSam-Conf
Single-Codec EnCodec-1VQ 3.36 ± 0.06 3.68 ± 0.08 28.7 0.669
Commitment Loss

EnCodec-4VQ 3.70 ± 0.10 3.86 ± 0.08 15.9 0.725


0.25
EnCodec-8VQ 3.77 ± 0.12 4.02 ± 0.11 11.4 0.738
TiCodec-1VQ 3.68 ± 0.08 3.97 ± 0.07 19.4 0.683
0.20 Single-Codec 4.02 ± 0.07 4.13 ± 0.09 13.2 0.792

0.15
3.7. Zero-shot TTS Evaluation

0.10 To evaluate the performance of codecs applied in speech synthe-


sis tasks, we train VALL-E [1] using discrete tokens extracted
0 50000 100000 150000 200000 250000 300000 from EnCodec in the number of 1,4,8 codebooks, TiCodec in
Step
Figure 2: The commitment loss of different codec while training. 1 codebook, and Single-Codec. We conduct naturalness Mean
Opinion Score (N-MOS) and speaker similarity MOS (S-MOS)
3.5. Commitment Loss Analysis for subjective evaluation of the synthesized speech. The test set
We further analyze the commitment loss in the training to ex- consists of 30 sentences, including Chinese and English speech.
plore the impact of different designed modules on the single- The 20 listeners who are Chinese Mandarin native speakers and
codebook codec. Commitment loss is the difference between familiar to English are invited to participate in each MOS test.
representations before and after quantization. The degree of Meanwhile, we calculate the word error rate (WER) using an
convergence of the commitment loss can reflect the relation- ASR model6 [33] to measure speech intelligibility. We also use
ship between the encoder output and the cluster center in the WeSpeaker [32] to extract speaker embedding to calculate the
codebook. As shown in Figure 2, the commitment loss of VQ- speaker embedding cosine similarity.
VAE tends to diverge after model training, indicating that the Table 2 shows the subjective and objective results. Single-
entanglement of time-invariant global information and time- Codec outperforms other models in terms of naturalness and
variant content information hinders forming a limited variety of speaker similarity. In single-codebook scenes, TiCodec-1VQ
content-related speech units. After considering time-invariant and Single-Codec are significantly better than other codec mod-
decoupled modelling, the loss of Ref-short increases slowly, els in speaker similarity, naturalness, and stability. This is
indicating the effectiveness of global information disentangle- because decoupling global information makes the frame-level
ment for speech unit learning. Ref-long further verifies this re- codebook pay more attention to content modelling and enables
sult, illustrating the effectiveness of a longer reference segment. more global information transmission. Meanwhile, Single-
The loss curve of Ref-HybSam is flat, indicating that the hy- Codec performs better than Ticodec, indicating the effective-
brid sampling module effectively improves codec performance. ness of continuous global representation and additional con-
Moreover, the losses of the models with context modelling via tent modelling. In addition, Single-Codec exceeds multiple-
the BLSTM module are all converged. It demonstrates that the codebook codecs regarding speaker similarity and naturalness
models have learned stable phonetic units before quantization, while WER is slightly higher than Encodec-8VQ. This is
indicating the effectiveness of context modelling in codecs. mainly because the higher bandwidth brings higher-resolution
Furthermore, considering the results presented in Table 1, speech unit perception.
we observe that the commitment loss is not strictly inversely
related to reconstruction quality. However, the convergence sta-
tus of the commitment loss (divergence, flat, convergence) is
4. Conclusion
indeed associated with reconstruction quality. Specifically, the In this paper, we propose Single-Codec, the first single-
converged codec surpasses the codec which is not converged. codebook codec dedicatedly designed for LLM-based speech
This result further highlights the significance of achieving a sta- generation. Single-Codec employs a disentangled VQ-VAE on
ble clustering center in the single codebook codec, which di- Mel Spectrograms to decouple speech into time-invariant global
rectly impacts the overall reconstruction quality. embedding and one phonetically-rich discrete sequence quan-
3.6. Speech Reconstruction Evaluation tized by one codebook. Furthermore, the encoder is enhanced
with a BLSTM module for contextual modelling, a hybrid sam-
We compare the performance in speech reconstruction of the
pling module to alleviate distortion from upsampling and down-
proposed Single-Codec with other codecs. The results, as pre-
sampling, and a resampling module to encourage discrete units
sented in Table 1, demonstrate that despite lower bandwidth,
to carry more phonetic-relevant information with lower short-
the proposed Single-Codec surpasses other codecs with 1 code-
time variance. In experiments, compared with multi-codebook
book and is on par with the TiCodec with 2 codebooks in
codecs, e.g. EnCodec and TiCodec, Single-Codec demonstrates
terms of reconstruction quality and speaker similarity. VQVAE
higher speech reconstruction quality with lower bandwidth of
performs better than EnCodec with 1 codebook, demonstrat-
only 304bps, and enables a higher-quality LLM-TTS with bet-
ing the high quantization efficiency of codecs operated on Mel
ter naturalness and intelligibility. In the future, we will focus on
Spectrogram. Compared to TiCodec, which also quantized the
developing a more efficient single-codebook codec for speech
decoupled time-invariant information, Single-Codec achieves
reconstruction and speech synthesis.
higher speaker similarity and reconstruction quality, indicating
the effectiveness of continuous time-invariant representations
and longer reference length. 6 [Link]
5. References [22] Y. Shi, H. Bu, X. Xu, S. Zhang, and M. Li, “AISHELL-3: A
multi-speaker mandarin TTS corpus and the baselines,” CoRR,
[1] C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, vol. abs/2010.11567, 2020.
Y. Liu, H. Wang, J. Li, L. He, S. Zhao, and F. Wei, “Neural codec
language models are zero-shot text to speech synthesizers,” CoRR, [23] A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu,
vol. abs/2301.02111, 2023. W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer:
Convolution-augmented transformer for speech recognition,” in
[2] E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin,
INTERSPEECH. ISCA, 2020, pp. 5036–5040.
O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour,
“Speak, read and prompt: High-fidelity text-to-speech with mini- [24] J. Chung, Çaglar Gülçehre, K. Cho, and Y. Bengio, “Empirical
mal supervision,” CoRR, vol. abs/2302.03540, 2023. evaluation of gated recurrent neural networks on sequence model-
[3] A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity ing,” ArXiv, vol. abs/1412.3555, 2014.
neural audio compression,” CoRR, vol. abs/2210.13438, 2022. [25] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for
[4] X. Zhang, D. Zhang, S. Li, Y. Zhou, and X. Qiu, “Speechtok- image recognition,” in CVPR. IEEE Computer Society, 2016,
enizer: Unified speech tokenizer for speech large language mod- pp. 770–778.
els,” CoRR, vol. abs/2308.16692, 2023. [26] D. Yang, S. Liu, R. Huang, J. Tian, C. Weng, and Y. Zou, “Hifi-
[5] S. Ji, M. Fang, Z. Jiang, R. Huang, J. Zuo, S. Wang, and codec: Group-residual vector quantization for high fidelity audio
Z. Zhao, “Language-codec: Reducing the gaps between discrete codec,” CoRR, vol. abs/2305.02765, 2023.
codec representation and speech language models,” CoRR, vol. [27] X. Zhang, L. Xue, Y. Wang, Y. Gu, X. Chen, Z. Fang, H. Chen,
abs/2402.12208, 2024. L. Zou, C. Wang, J. Han, K. Chen, H. Li, and Z. Wu, “Amphion:
[6] Y. Ren, T. Wang, J. Yi, L. Xu, J. Tao, C. Zhang, and J. Zhou, An open-source audio, music and speech generation toolkit,”
“Fewer-token neural speech codec with time-invariant codes,” CoRR, vol. abs/2312.09911, 2023.
CoRR, vol. abs/2310.00014, 2023. [28] C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A short-
[7] Z. Borsos, M. Sharifi, D. Vincent, E. Kharitonov, N. Zeghidour, time objective intelligibility measure for time-frequency weighted
and M. Tagliasacchi, “Soundstorm: Efficient parallel audio gen- noisy speech,” in ICASSP. IEEE, 2010, pp. 4214–4217.
eration,” CoRR, vol. abs/2305.09636, 2023. [29] I.-T. Recommendation, “Perceptual evaluation of speech quality
[8] D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, X. Chang, J. Shi, (pesq): An objective method for end-to-end speech quality as-
S. Zhao, J. Bian, X. Wu, Z. Zhao, S. Watanabe, and H. Meng, sessment of narrow-band telephone networks and speech codecs,”
“Uniaudio: An audio foundation model toward universal audio Rec. ITU-T P. 862, 2001.
generation,” CoRR, vol. abs/2310.00704, 2023. [30] R. F. Kubichek, “Mel-cepstral distance measure for objective
[9] A. Ziv, I. Gat, G. L. Lan, T. Remez, F. Kreuk, A. Défossez, speech quality assessment,” Proceedings of IEEE Pacific Rim
J. Copet, G. Synnaeve, and Y. Adi, “Masked audio genera- Conference on Communications Computers and Signal Process-
tion using a single non-autoregressive transformer,” CoRR, vol. ing, vol. 1, pp. 125–128 vol.1, 1993.
abs/2401.04577, 2024.
[31] T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and
[10] J. Betker, “Better speech synthesis through scaling,” CoRR, vol. H. Saruwatari, “UTMOS: utokyo-sarulab system for voicemos
abs/2305.07243, 2023. challenge 2022,” in INTERSPEECH. ISCA, 2022, pp. 4521–
[11] A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural dis- 4525.
crete representation learning,” ArXiv, vol. abs/1711.00937, 2017. [32] H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang,
[12] P. Esser, R. Rombach, and B. Ommer, “Taming transformers for Y. Deng, and Y. Qian, “Wespeaker: A research and production
high-resolution image synthesis,” in CVPR. Computer Vision oriented speaker embedding learning toolkit,” in ICASSP. IEEE,
Foundation / IEEE, 2021, pp. 12 873–12 883. 2023, pp. 1–5.
[13] S. Lee, W. Ping, B. Ginsburg, B. Catanzaro, and S. Yoon, “Bigv- [33] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and
gan: A universal neural vocoder with large-scale training,” in I. Sutskever, “Robust speech recognition via large-scale weak su-
ICLR. [Link], 2023. pervision,” in ICML, ser. Proceedings of Machine Learning Re-
search, vol. 202. PMLR, 2023, pp. 28 492–28 518.
[14] M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural
networks,” IEEE Trans. Signal Process., vol. 45, pp. 2673–2681,
1997.
[15] S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
Neural Comput., vol. 9, no. 8, pp. 1735–1780, 1997.
[16] E. Vincent, R. Gribonval, and C. Févotte, “Performance measure-
ment in blind audio source separation,” IEEE Trans. Speech Audio
Process., vol. 14, no. 4, pp. 1462–1469, 2006.
[17] J. Shi, H. Inaguma, X. Ma, I. Kulikov, and A. Y. Sun, “Multi-
resolution hubert: Multi-resolution speech self-supervised learn-
ing with masked unit prediction,” CoRR, vol. abs/2310.02720,
2023.
[18] H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen,
and Y. Wu, “Libritts: A corpus derived from librispeech for text-
to-speech,” in INTERSPEECH. ISCA, 2019, pp. 1526–1530.
[19] E. Bakhturina, V. Lavrukhin, B. Ginsburg, and Y. Zhang, “Hi-fi
multi-speaker english TTS dataset,” in Interspeech. ISCA, 2021,
pp. 2776–2780.
[20] Z. Liu and B. K.-W. Mak, “Cross-lingual multi-speaker text-to-
speech synthesis for voice cloning without using parallel corpus
for unseen speakers,” arXiv: Audio and Speech Processing, 2019.
[21] H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “AISHELL-1: an
open-source mandarin speech corpus and a speech recognition
baseline,” in O-COCOSDA. IEEE, 2017, pp. 1–5.

Common questions

Powered by AI

The Single-Codec achieves a balance between lower bandwidth usage and high-quality speech reconstruction by employing a single-codebook design that effectively disentangles time-invariant features and encodes speech into a phonetically rich discrete sequence using VQ-VAE. Its advanced modules for contextual modeling (BLSTM), hybrid sampling, and resampling enhance feature extraction and clustering, allowing high-quality reconstruction at just 304bps .

The BLSTM module in the Single-Codec architecture is used for contextual modeling, which helps in discovering correlations between adjacent frames. This enhances the clustering efficiency of speech content, facilitating the formation of stable phonetic units prior to quantization and contributing to the overall speech reconstruction quality .

The Single-Codec outperforms other multi-codebook codecs in speech reconstruction quality, even at lower bandwidth, because it uses a novel approach to feature disentanglement, focusing on continuous time-invariant representations and longer reference segments. These techniques enhance phonetic detail capture and transmission, leading to higher intelligibility and speaker similarity in the synthesized speech even in zero-shot scenarios .

Subjective evaluations, measured through MOS scores, correlate well with objective metrics like STOI and PESQ in assessing the Single-Codec's performance. It achieves high naturalness and speaker similarity scores in MOS tests, paralleling its superior objective metrics performance. This indicates that the codec effectively balances technical quality with perceived sound quality, reinforcing its advancement over other codecs .

Single-Codec contributes to improved zero-shot text-to-speech synthesis performance by employing disentangled VQ-VAE on Mel Spectrograms, thus efficient at transmitting phonetics-relevant information with lower bandwidth. This enhances naturalness and speaker similarity, even surpassing TiCodec-2VQ, suggesting effective global information decoupling which ensures that frame-level codebooks emphasize content modeling .

The resampling module in the Single-Codec system plays a crucial role by encouraging the extraction of more phonetics-relevant information with lower short-time variance. This is achieved through a bottleneck operation involving downsampling and residual connection after upsampling, which helps stabilize the learning of acoustic unit representations, thereby enhancing the codec’s performance .

The key objective metrics used to evaluate the performance of codecs include Short-Time Objective Intelligibility (STOI), Perceptual Evaluation of Speech Quality (PESQ), Mel-Cepstral Distortion (MCD), and MOS-based speaker similarity (SPK). Single-Codec shows superior performance, achieving higher STOI and PESQ scores, and lower MCD compared to other codecs like EnCodec and TiCodec, even at reduced bandwidth, indicating its efficient design and enhanced speech reconstruction capabilities .

The primary purpose of introducing a global encoder in the Single-Codec architecture is to decouple time-invariant features from speech data. This decoupling enables the codec to focus more on phonetic information while embedding it into discrete units, thus improving the efficiency and quality of speech reconstruction by reducing the amount of frame-level encoding needed .

The hybrid sampling module enhances the Single-Codec's performance by combining convolution and pooling for downsampling, and transposed convolution and replication for upsampling. This combination helps alleviate distortion during the upsampling and downsampling processes, ensuring more accurate reconstruction of speech information .

Commitment loss analysis is significant as it reflects the relationship between encoder output and the cluster center in the codebook, indicating how well the model learns content-related speech units. In the case of Single-Codec, a converged commitment loss signifies stable and effective learning of phonetic units, enabling better reconstruction quality. This analysis showcases the impact of modules like the BLSTM and hybrid sampling in stabilizing and improving the codec's performance .

You might also like