Spectral subtraction and RASTA-filtering in text-dependent HMM-based speaker verification
Daniel Hardt, Klaus Fellbaum · 2002
In real text-dependent telephone-based speaker verification systems, both, additive and convolutional noise influence the error rate considerably. In this paper, different procedures which make a speaker verification system more robust against noise are compared. We either use spectral subtraction in addition to MFCC-feature extraction or only PLP and RASTA-PLP (without spectral subtraction). Considering spectral subtraction two modifications were examined: one version which was preconnected to the system and a second one being integrated into the MFCC computation. The first version has the advantage that the window length can be chosen independently of those of the MFCC procedure. This led to better results. However, the most effective procedure for telephone speech data is the J-RASTA-PLP, but the estimation of the optimal J factor is difficult. At first we used a fixed J factor based on off-line measurement of noise power. Finally, we performed some experiments to optimize the system with the adaptive estimation of the J factor during utterance. This procedure is based on the method of spectral mapping which has been shown to be very effective in automatic speech recognition.