A dynamic automatic noisy speech recognition (DANSR) system for a single-channel hybrid noisy industrial environment
Sheuli Paul, Michael M. Richter · Proceedings of meetings on acoustics · 2013
A dynamic noisy speech recognition system is developed to recognize single-channel speaker independent small spoken commands in a hybrid noisy industrial environment. The hybrid noise is environmental mixed noise distinguished as: (i) Strong, (ii) Time varying steady-unsteady, (iii) Mild. In [1] we presented a principal architecture for solving this problem. There were, however, several parts in the solution system missing. They are mainly on the technical level of speech recognition. These will be presented in this paper. A new adaptive feature extraction technique based on local trigonometric transformation (LTT) is introduced and examined. This is adapted with psychoacoustic quantities such as : a) Bark scaled critical band spectrum, b) loudness scale, and c) perceptual entropy. Here the spectral analysis is done by rising cut-off function, folding operation and discrete cosine transformation (DCT-IV) instead of Fourier transform. Then inverse DCT-IV and unfolding operation result in perceptual LTT (PLTT) features. These are classified by Gaussian mixture model (GMM) and recognized by hidden Markov model (HMM). The new PLTT features are more efficient and perceptually meaningful than the standard feature extraction techniques. The DANSR system is a novel solution for small commands to a long existing hybrid noise problem.