Improved ASR System based on CNN-LSTM Acoustic Model for Mobile Voice

Belabbas Soumeya, Djamel Addou, Selouani Sid Ahmed · 2023

The increasing demand for data communication is not only limited to wired connections but has also spread to the wireless world. People now desire access to information while on the move, and this has become possible with the deployment of advanced technologies. Furthermore, mobile devices with improved speech input interfaces provide users with easy access to these data services. In automatic mobile speech recognition systems, the variability caused by the environment (presence of noise) leads to different types of disparities between observations and acoustic models, which generates low recognition rates in noisy conditions. The present work aims to improve the performance of the automatic speech recognition system “ASR” by using more robust acoustic parameters such as Mel spectrogram features, Mel Frequency Cepstral Coefficients “MFCC”, Power Normalized Cepstral Coefficients “PNCC”, Gammatone Cepstral coefficients “GTCC” and the robust MFCC extracted from the “ETSI-AFE” standard “The European Telecommunications Standards Institute advanced frontend”. This parameter “AFE-MFCC” based on speech enhancement yields higher recognition rates in noisy environments. To demonstrate the robustness and strength of this parameter, we integrated a robust acoustic model resulting from the concatenation of two neuronal models, the convolutional neural network “CNN” and the Long Short-Term Memory “LSTM” network. This combined system “CNN-LSTM” based on the “AFE-MFCC” parameter leads to better recognition rates in low signal-to-noise ratio “SNR” environments.

Read the paper · More papers on PaperTik