Robust Features in Deep Neural Networks for Transcoded Speech Recognition DSR and AMR-NB

Lallouani Bouchakour, Mohamed Debyeche, Ahmed Krobba · 2024

Automatic Speech Recognition (ASR) performance in mobile communications degrades significantly if the environment includes many sources of variability, such as when the test environment differs from the training environment and when the acoustical environment includes disturbances like noise, channel distortion, speaker differences, and mobile codecs. In this work, we have used two architectures for speech recognition in mobile networks. The first one is Distributed Speech Recognition based on DSR coding, and the second architecture is based on Adaptive Multi-Rate Narrow-Band (AMR-NB). We propose a novel robust feature extraction (Front-End) technique to improve speech recognition performance in noisy mobile communications. This technique utilizes special parameters such as Gabor features, Power Normalized Spectrum Gabor filter (PNS-Gabor), and Power Standardized Cepstral Coefficients (PNCC). These features take into account psychoacoustic effects like the temporal masking effect and have different distributions of filter banks and filter forms to better model human perception. In the back end, we investigated speech classification systems using Continuous Hidden Markov Models (CHMM) and Deep Neural Networks (DNN). Based on the results obtained in noisy mobile communications, the proposed features PNS-Gabor and PNCC show significant improvements over conventional acoustic features such as Mel frequency cepstral coefficients (MFCC).

Read the paper · More papers on PaperTik