Improved Far-field Speaker Recognition Method Based Geometry Acoustic Simulation and SpecAugment

Yingzi Lian, Jing Pang · 2021

Deep learning method has made a great breakthrough in the field of speaker recognition. However, the performance of deep learning method based neural network often degrade in far-field speaker recognition task, with respect to the near-field. One of the indispensable factors is that deep learning performance has a great impact on the quality and diversity of data. Utterance in far-field scenario always mixed with variant reverberation and noise which could interfere the prediction of model. In this paper, we proposed a novel method comprised of geometry acoustic simulation (GAS) and SpecAugment, which makes the performance of deep learning method on far-field speaker recognition task more stable and excellent. Based on near-field speech dataset, by means of physical simulation, far-field speech with different types and characters can be generated to provide as many scenes as possible for speaker recognition. Simultaneously, the SpecAugment method has been applied to enhance the diversity of utterance data on mel-spectrogram. Then, ResNet34 and AM-Softmax Loss were used as the most commonly used methods for feature extraction and model training. Our proposed method has been trained on Voxceleb1 and V oxceleb2 dataset, and test on Vox celeb original test set. Evaluations show that our proposed method has 46% and 40% of EER and minDCF relatively improvement on far-field speaker recognition task.

Read the paper · More papers on PaperTik