Reshape Dimensions Network for Speaker Recognition
Reshape Dimensions Network for Speaker Recognition
Ivan Yakovlev, Rostislav Makarov, Andrei Balykin, Pavel Malov, Anton Okhotnikov,
Nikita Torgashov
Abstract
In this paper, we present Reshape Dimensions Network
(ReDimNet), a novel neural network architecture for extract-
arXiv:2407.18223v2 [[Link]] 25 Sep 2024
1. Introduction
Speaker recognition is a specialized field aiming at identifying
or verifying individuals through their distinct voice features.
In this domain, deep neural networks have emerged as a ma-
jor technology for extracting speaker embeddings that are used
for multiple tasks including Speaker Verification (SV), Speaker
Identification, Speaker Diarization, and others. Extensive re-
search has been conducted in the SV area, which includes the Fig. 1 Computational Cost vs. Average Equal Error Rate.
development of new datasets [1–4], model architecture design- EER is averaged over three Voxceleb1 protocols: Vox1-O,
ing [5–18], and inventing new loss functions [19, 20]. Vox1-E, Vox1-H. The model size is shown by the area of a cir-
A variety of architectures have emerged including 1D [5–7, cle, model family is indicated by a color. Complexity is as-
9, 10] and 2D [14–18] convolutional neural networks (CNNs), sessed using thop library with an input signal of 2 seconds. A
their hybrids that incorporate 2D CNN stem before 1D TDNN- short dashed line represents scaling the ReDimNet architecture
like backbone [8,11,13], as well as self-attention networks [12]. using the Additive Angular Margin loss function [19], dashed
Each architectural approach brings its unique set of advantages line - using the SphereFace2 loss [20].
with 1D models offering efficiency and direct temporal analy-
sis, 2D architectures providing frequency translational invari-
ance [21], and hybrid systems aiming to deliver the best of both ing optimal performance under varying computational resource
worlds. Additionally, design approaches can be split into macro constraints. Our experimental results demonstrate that ReDim-
and micro designs, with micro designs involving modifications Net outperforms many other architectures and achieves state-
like substituting traditional 1D ResBlocks with Res2Net blocks of-the-art performance on public benchmarks while reducing
within the ECAPA-TDNN architecture [6], and macro designs inference time and model size.
incorporating a 2D stem ahead of TDNN-like models [8,11,13]
leading to a two-stage architecture that transitions 2D −
→ 1D. 2. Model Architecture
In this paper, we introduce ReDimNet1 , a novel neural net-
work architecture based on the dimensionality reshaping of fea- In this section, we detail the design of the proposed architecture
ture maps between 2D and 1D representations, enabling seam- influenced by two main concepts. Firstly, to leverage the bene-
less integration of 1D and 2D blocks. ReDimNet exhibits scal- fits of residual connections, we incorporate them extensively in
ability across various model sizes, while consistently achiev- ReDimNet. Secondly, based on the success of models utilizing
both 1D and 2D blocks for speech processing and SV, our ar-
1 [Link] chitecture integrates both types of blocks to boost performance.
Fig. 2 ReDimNet architecture scheme. Digits 1,2,3 and 4 describe the order of operators and blocks execution in a single model stage,
where C - number of channels, F - number of frequency bins, T - number of timestamps.
1D block
block [30] with fwSE [21]. As 1D blocks we used same (a) 1D 1D Conv 0.65 0.85 1.54 1.01
version of ConvNeXt-like block with or inplace of (c) Trans- MHA 0.69 0.82 1.45 0.99
1D Conv + MHA 0.59 0.79 1.47 0.95
former block [31]. ConvNext block 0.68 0.83 1.46 0.99
2D block fwse-ResNet block 0.64 0.82 1.48 0.98
ResNet block 0.61 0.80 1.48 0.96
applied finetuning on longer utterances with some augmenta-
tions turned off and tweaked the parameters of a loss function.
This second training stage is well-known as Large-Margin (LM) 4.2. Ablation studies
finentuning strategy [24]. All models were trained using the We also conducted a thorough study of how different compo-
wespeaker [25] training pipeline. nents of ReDimNet architecture affect its performance. This
research includes studying the role of 1D and 2D blocks for
3.1. Pretraining stage speech signal processing, assessing the impact of different loss
functions, and optimizing group sizes and steps in convolutions
For pretraining, we used a default voxceleb2 recipe from for accuracy and efficiency improvement. All ablation studies
wespeaker pipeline with minor adjustments. 2-second seg- were carried out on the ReDimNet-B2 architecture.
ments were selected randomly from each signal, and various
augmentations with MUSAN dataset [26] (noise, music, bab- Table 3 Ablation study on loss function configuration (EER,%)
ble) alongside the RIR dataset [27] were applied following the
augmentation recipe from [14]. A two-fold speed augmenta- Loss Type Vox1-O Vox1-E Vox1-H Average
tion [28], with factors of 0.9 and 1.1, was employed to generate AAM-SC 0.57 0.91 1.60 1.03
additional speakers within the training dataset. In this stage, the AAM 0.68 0.83 1.46 0.99
AAM-softmax margin penalty was scheduled as follows: first SF2-A 0.63 0.80 1.39 0.94
20 epochs it was kept at 0.0, then for the next 20 epochs it ex- SF2-C 0.57 0.76 1.32 0.88
ponentially rose to 0.2 and then was kept constant till the end
of training. We used Exponential Decay with Warmup learn-
ing rate scheduler with 6 epochs warmup, lrmax = 1e−1 and 4.2.1. Block types
lrmin = 1e−5 . In order to identify optimal configurations of ReDimNet archi-
tecture, we compared three types of 2D-blocks: basic ResNet
3.2. Large-Margin Finetuning stage block, basic ResNet FWSE block, and a ConvNext block.
While minimal differences were observed, basic ResNet block,
At the finetuning stage [24], AAM-softmax margin was set to however, slightly outperformed others by a small margin (see
constant 0.5 value, with length of training utterances expanded Table 2).
to 6 seconds. Speed perturbations were turned off during this Our further analysis was focused on the 1D block type,
stage. where we assessed a range of options including sequences of 1D
convolutional ConvNeXt-like blocks (Fig. 3), multi-head atten-
3.3. Evaluation tion (MHA) (Fig. 3), FC layers, skip connections, and a hybrid
of 1D convolutional blocks with MHA (1D Conv + MHA). Skip
The performance of models is assessed using cleaned protocols connections appeared to be the least effective approach, which
of VoxCeleb1 [32] test set, employing the Equal Error Rate underscored the importance of the 1D block within the ReD-
(EER) and the minimum Detection Cost Function (minDCF) imNet architecture. FC layers performed slightly better, sug-
with Ptarget = 0.01 and CF A = CM iss = 1. We scored each gesting the importance of a temporal context. 1D convolutional
model with cosine backend utilizing full utterance length as in- and MHA blocks have proven to be the most efficient configura-
put and additionally applied a top-300 adaptive s-normalization tions, and a combination of MHA and 1D convolutional blocks
(AS-Norm) [33] of cosine scores (see Table 4). delivered the best performance (see Table 2).
Table 4 Evaluation results on the VoxCeleb1-Cleaned protocols without QMFs. For the report, we calculated the equal error rate
(EER) and the minimum detection cost function (minDCF). GMACs were measured on 2-s long segments. * - means values have been
estimated. Open source models from the WeSpeaker or ECAPA2 repositories were retested in our environment.
4.2.2. Loss studies Table 5 Evaluation results on Speakers In The Wild core-core
protocol [35], VOiCES from a Distance Challenge Evaluation
Furthermore, we explored the effectiveness of various loss func- Set [36] and VoxCeleb1-B protocol [37] (EER, %).
tions (see Table 3). Specifically, we evaluated SphereFace
losses (SF2) with A and C configurations [20], Additive Angu- Model SITW VOiCES Vox1-B Average
lar Margin Loss (AAM), and Additive Angular Margin loss with CAM++ 1.34 6.30 2.79 3.48
SubCenters (AAM-SC) [19]. Based on the testing results, we ECAPA (C=1024) 1.67 5.31 3.48 3.49
found SphereFace type C to be the most effective loss function ResNet293 1.67 5.14 2.23 3.01
providing the largest performance improvement in the bench- ECAPA2 3.64 13.26 1.81 6.24
marks. ReDimNet-B6 0.77 3.19 1.66 1.87
5. Results 6. Conclusions
Testing results of all proposed ReDimNet architecture config- In this paper we introduced ReDimNet - a novel neural net-
urations are presented in Table 4. We compared ReDimNet work architecture designed for the extraction of utterance-level
on the VoxCeleb1 protocols with publicly available models and speaker representations. It combines dimensionality reshap-
grouped them based on the number of parameters and multiply- ing, dynamic transitions between 1D and 2D representations,
accumulate operations (MACs) for comparison purposes. and 2D and 1D blocks. Through a comprehensive evaluation,
ReDimNet demonstrated:
In particular, our ReDimNet-B1 model achieves compara-
ble results to NeXt-TDNN [7] on the Vox1-H protocol, but has • architecture adaptability and scalability across multiple
a slightly larger number of parameters and MACs. ReDimNet- configurations;
B3 outperforms Gemini DF-ResNet60 [18] and ECAPA (C = • top balance between computational efficiency and perfor-
1024) with an advantage in model size. ReDimNet-B5 further mance;
improves upon the B3 version, consistently achieving the low- • strong results on the VoxCeleb1-H (cleaned) protocol, with
est EER and minDCF, compared to DF-ResNet233 [17], which an Equal Error Rate (EER) of 1.00%;
has the similar number of parameters and MACs. Moreover, our • advanced generalization ability on out-of-domain test sets.
largest model, ReDimNet-B6, delivers even better results while
In summary, ReDimNet architecture achieves competitive
having significantly fewer parameters and MACs than ECAPA
performance on all tests compared to other state-of-the-art
2 [13] and ResNet293 [25, 30].
speaker recognition models, while also offering favorable com-
Furthermore, we subjected the best models of various archi- putational efficiency. Its adaptability and superior performance
tectures to additional out-of-domain testing (see Table 5). These make it a valuable contribution to the speaker recognition field
results demonstrate that ReDimNet-B6 outperforms other archi- and a promising solution for real-world applications.
tectures with a significant gap on unseen data domains.
7. References [21] J. Thienpondt, B. Desplanques, and K. Demuynck, “Integrating
frequency translational invariance in tdnns and frequency posi-
[1] J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep tional information in 2d resnets to enhance speaker verification,”
speaker recognition,” in Interspeech 2018. ISCA, Sep. 2018. arXiv preprint arXiv:2104.02370, 2021.
[2] Y. Lin, X. Qin, G. Zhao, M. Cheng, N. Jiang, H. Wu, and M. Li, [22] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan,
“Voxblink: A large scale speaker verification dataset on camera,” T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al., “Pytorch:
2023. An imperative style, high-performance deep learning library,” Ad-
[3] S. Zheng, L. Cheng, Y. Chen, H. Wang, and Q. Chen, “3d-speaker: vances in neural information processing systems, vol. 32, 2019.
A large-scale multi-device, multi-distance, and multi-dialect cor- [23] K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statis-
pus for speech representation disentanglement,” 2023. tics pooling for deep speaker embedding,” in Interspeech 2018.
[4] I. Yakovlev, A. Okhotnikov, N. Torgashov, R. Makarov, Y. Vo- ISCA, Sep. 2018.
evodin, and K. Simonchik, “VoxTube: a multilingual speaker [24] Y. Liu, L. He, and J. Liu, “Large margin softmax loss for speaker
recognition dataset,” in Proc. INTERSPEECH 2023, 2023, pp. verification,” arXiv preprint arXiv:1904.03479, 2019.
2238–2242.
[25] H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang,
[5] D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudan- Y. Deng, and Y. Qian, “Wespeaker: A research and production
pur, “X-vectors: Robust dnn embeddings for speaker recognition,” oriented speaker embedding learning toolkit,” in ICASSP 2023-
in 2018 IEEE International Conference on Acoustics, Speech and 2023 IEEE International Conference on Acoustics, Speech and
Signal Processing (ICASSP), 2018, pp. 5329–5333. Signal Processing (ICASSP). IEEE, 2023, pp. 1–5.
[6] B. Desplanques, J. Thienpondt, and K. Demuynck, “Ecapa- [26] D. Snyder, G. Chen, and D. Povey, “Musan: A music, speech, and
tdnn: Emphasized channel attention, propagation and ag- noise corpus,” arXiv preprint arXiv:1510.08484, 2015.
gregation in tdnn based speaker verification,” arXiv preprint
arXiv:2005.07143, 2020. [27] I. Szöke, M. Skácel, L. Mošner, J. Paliesek, and J. Černockỳ,
“Building and evaluation of a real room impulse response
[7] H.-J. Heo, U.-H. Shin, R. Lee, Y. Cheon, and H.-M. Park, “Next- dataset,” IEEE Journal of Selected Topics in Signal Processing,
tdnn: Modernizing multi-scale temporal convolution backbone vol. 13, no. 4, pp. 863–876, 2019.
for speaker verification,” 2023.
[28] T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmen-
[8] T. Liu, R. K. Das, K. A. Lee, and H. Li, “Mfa: Tdnn with multi- tation for speech recognition.” in Interspeech, vol. 2015, 2015, p.
scale frequency-channel attention for text-independent speaker 3586.
verification with short utterances,” 2022.
[29] Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie,
[9] Z. Zhao, Z. Li, W. Wang, and P. Zhang, “Pcf: Ecapa-tdnn with “A convnet for the 2020s,” 2022.
progressive channel fusion for speaker verification,” 2023.
[30] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning
[10] Y.-Q. Yu and W.-J. Li, “Densely connected time delay neural net- for image recognition,” in 2016 IEEE Conference on Computer
work for speaker verification,” in Annual Conference of the Inter- Vision and Pattern Recognition (CVPR), 2016, pp. 770–778.
national Speech Communication Association (INTERSPEECH),
2020, pp. 921–925. [31] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N.
Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,”
[11] H. Wang, S. Zheng, Y. Chen, L. Cheng, and Q. Chen, “CAM++: 2023.
A Fast and Efficient Network for Speaker Verification Using
Context-Aware Masking,” in Proc. INTERSPEECH 2023, 2023, [32] A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, “Voxceleb:
pp. 5301–5305. Large-scale speaker verification in the wild,” Computer Science
and Language, 2019.
[12] Y. Zhang, Z. Lv, H. Wu, S. Zhang, P. Hu, Z. Wu, H. yi Lee, and
H. Meng, “Mfa-conformer: Multi-scale feature aggregation con- [33] S.-C. Yin, R. Rose, and P. Kenny, “Adaptive score normalization
former for automatic speaker verification,” 2022. for progressive model adaptation in text independent speaker ver-
ification,” in 2008 IEEE International Conference on Acoustics,
[13] J. Thienpondt and K. Demuynck, “Ecapa2: A hybrid neural net- Speech and Signal Processing, 2008, pp. 4857–4860.
work architecture and training strategy for robust speaker embed-
dings,” 2024. [34] M. Tan and Q. V. Le, “Efficientnet: Rethinking model scaling for
convolutional neural networks,” 2020.
[14] D. Garcia-Romero, G. Sell, and A. Mccree, “Magneto: X-vector
magnitude estimation network plus offset for improved speaker [35] M. McLaren, L. Ferrer, D. Castan, and A. Lawson, “The speakers
recognition.” in Odyssey, 2020, pp. 1–8. in the wild (sitw) speaker recognition database.” in Interspeech,
2016, pp. 818–822.
[15] T. Zhou, Y. Zhao, and J. Wu, “Resnext and res2net structures for
speaker verification,” 2020. [36] M. K. Nandwana, J. van Hout, M. McLaren, C. Richey, A. Law-
son, and M. A. Barrios, “The voices from a distance challenge
[16] Y. Chen, S. Zheng, H. Wang, L. Cheng, Q. Chen, and J. Qi, 2019 evaluation plan,” 2019.
“An Enhanced Res2Net with Local and Global Feature Fusion for
Speaker Verification,” in Proc. INTERSPEECH 2023, 2023, pp. [37] K. Nam, Y. Kim, J. Huh, H. S. Heo, J. weon Jung, and J. S. Chung,
2228–2232. “Disentangled representation learning for multilingual speaker
recognition,” 2023.
[17] B. Liu, Z. Chen, S. Wang, H. Wang, B. Han, and Y. Qian, “DF-
ResNet: Boosting Speaker Verification Performance with Depth-
First Design,” in Proc. Interspeech 2022, 2022, pp. 296–300.
[18] T. Liu, K. A. Lee, Q. Wang, and H. Li, “Golden gemini is all you
need: Finding the sweet spots for speaker verification,” 2023.
[19] J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive
angular margin loss for deep face recognition,” in Proceedings of
the IEEE/CVF conference on computer vision and pattern recog-
nition, 2019, pp. 4690–4699.
[20] B. Han, Z. Chen, and Y. Qian, “Exploring binary classification
loss for speaker verification,” in ICASSP 2023 - 2023 IEEE Inter-
national Conference on Acoustics, Speech and Signal Processing
(ICASSP), 2023, pp. 1–5.