A dynamic automatic noisy speech recognition system for a single-channel hybrid noisy industrial environment

Sheuli Paul · The Journal of the Acoustical Society of America · 2013

A dynamic noisy speech recognition system is developed to recognize single-channel small spoken commands in a hybrid noisy industrial environment. This hybrid system has three parts: (a) hybrid pre-processing to enhance noisy speech, (b) feature extraction for perceptual speech features, (c) classification and recognition for the DANSR's result. Here, the single-channel is only one microphone, and the hybrid noise is environmental mixed noise distinguished as: (i) strong, (ii) time varying steady-unsteady, and (iii) mild. A new adaptive feature extraction technique based on local trigonometric transformation (LTT) is introduced and examined. This is adapted with psychoacoustic quantities such as Bark scaled critical band spectrum, loudness scale, and perceptual entropy. Here the spectral analysis is done by rising cut-off function, folding operation, and discrete cosine transformation (DCT-IV) instead of Fourier transform. Then, inverse DCT-IV and unfolding operation result in perceptual LTT (PLTT) features. These are recognized by hidden Markov model (HMM). The new PLTT features are more efficient and perceptually meaningful than the standard feature extraction techniques. The DANSR system is a novel solution for small commands to a long existing hybrid noise problem.

Read the paper · More papers on PaperTik