Robust speaker recognition using spectro-temporal autoregressive models

Sri Harish Mallidi, Sriram Ganapathy, Hynek Heřmanský · 2013

Speaker recognition in noisy environments is challenging when there is a mis-match in the data used for enrollment and veri-fication. In this paper, we propose a robust feature extraction scheme based on spectro-temporal modulation filtering using two-dimensional (2-D) autoregressive (AR) models. The first step is the AR modeling of the sub-band temporal envelopes by the application of the linear prediction on the sub-band dis-crete cosine transform (DCT) components. These sub-band en-velopes are stacked together and used for a second AR mod-eling step. The spectral envelope across the sub-bands is ap-proximated in this AR model and cepstral features are derived which are used for speaker recognition. The use of AR models emphasizes the focus on the high energy regions which are rel-atively well preserved in the presence of noise. The degree of modulation filtering is controlled using AR model order param-eter. Experiments are performed using noisy versions of NIST 2010 speaker recognition evaluation (SRE) data with a state-of-art speaker recognition system. In these experiments, the proposed features provide significant improvements compared to baseline features (relative improvements of 20 % in terms of equal error rate (EER) and 35 % in terms of miss rate at 10 % false alarm).

Read the paper · More papers on PaperTik