A weighted denoising auto-encoder applied to Mel sub-bands for robust speech recognition
Faezeh Baniardalan, Ahmad Akbari, Babak Nasersharif · 2017
Sub-band speech processing is well-known in robust speech recognition. On the other hand, in recent years, deep neural networks have been widely used in speech recognition for acoustic modeling and also feature extraction and transformation. In this paper, we propose to consider Mel sub-band speech processing in denoising auto-encoder(DAE) training to benefit from both mentioned method properties. In this way, in the training process, we assign lower weights to the Mel subbands containing higher level of noise, while we assign higher weights to sub-bands including lower level of noise. Furthermore, we use restricted Boltzmann Machine for pre-training of DAE. Thus, DAE can train the noise behavior and performs better in noise reduction. Experimental results on Aorura2 database, show the proposed DAE improve performance of logarithm of Mel filter bank energies (LMFB) for noisy speech recognition about 47.8% in the best case.