Deep Speaker Embedding Using Hybrid Network of Multi-Feature Aggregation and Multi-Loss Fusion for TI-SV

Xiao Li, Xiao Hu, Xiao Chen, Hang Pan, Kun Niu · 2022 26th International Conference on Pattern Recognition (ICPR) · 2022

Text-independent speaker verification (TI-SV) refers to the process of verifying an individual’s claimed identity from a given speech utterance with unfixed content. Most deep speaker embedding networks of TI-SV apply temporal pooling or similar techniques for frame-level feature aggregation, and adopt a single loss function for training. In this paper, we propose a powerful hybrid network, named HN-MFML, which consists of a backbone and two sub-networks: speaker embedding extraction and speaker classification. The hybrid network not only incorporates global and local features, but also assigns adaptive weights to them. It adopts a modified ResNet-50 as backbone, using speaker embedding extraction sub-network to aggregate global and local features adaptively in frequency-time domain, which can be trained end-to-end by a loss. In addition, we add a speaker classification sub-network with another loss and explore a multi-loss fusion to jointly train for improving generalization. We demonstrate that our multi-feature aggregation and multi-loss fusion are superior for obtaining discriminative utterance-level embedding descriptors. We also show that HN-MFML achieves state-of-the-art performance by a significant margin compared with previous methods.

Read the paper · More papers on PaperTik