A speaker recognition method based on Res-SE-MFA and SpecAugment data enhancement

Hongbin He, Zhidan Li, Jixiang Cheng · 2022

Speaker recognition is a technology that uses voice signals to extract vocal features to distinguish the identity of different individuals and is widely used in mobile banking, security, and computer access control. In recent years, with the continuous development of deep learning technology, speaker recognition technology based on deep learning has also been developed rapidly and received wide attention. In this paper, we propose a speaker recognition method based on the Residual-Squeeze-Excitation Multilayer feature aggregation (Res-SE-MFA) framework to address the problems of weak generalization and poor robustness of deep learning on small sample datasets and further enhance the robustness of the system by using the SpecAugment data enhancement method on this basis. Experiments show that the comparison with ResNet34 improves by 10.01 % on EER and 2.99% on minDCF.

Read the paper · More papers on PaperTik