LSTM Based End-to-End Text-Independent Speaker Verification Using Raw Waveform

Jing Feng He, Pengwei Zhang, Liangjin Zhu · 2020

Speaker can be discriminated either at voice source level or vocal tract system level. Conventionally Mel-Frequency Cesptral Coefficients (MFCCs) or Mel filterbank energies are employed as input acoustic feature in neural network based speaker verification systems. In this paper, we investigate the LSTM based speaker verification using raw waveform as input feature. The basic LSTM based SV model and the model with attention layer are trained and optimized on two datasets using raw waveform feature and Fbank feature respectively. And experimental results show that compared with the model trained using Fbank feature, the model trained using raw waveform can achieve promising performance, raw waveform is a competitive acoustic feature for LSTM based speaker verification.

Read the paper · More papers on PaperTik