Performance Evaluation of Recurrent Neural Networks-LSTM and GRU for Automatic Speech Recognition

Dhiraj Kumar, Shahid Aziz · 2023

The work presented in this paper evaluates and compares the performance of the conventional Recurrent neural network(RNN), with elegant variations in the RNN architecture namely the LSTM and Gated Recurrent Units(GRU) for speech recognition. The conventional RNNs inept handling of the long-term dependencies due to the vanishing/exploding gradient problem led to the LSTMs and GRUs which specialises in handling long-term dependencies and does well in speech recognition tasks. The LSTM and GRU networks provide much greater frame precision than the RNNs while converging much more quickly. Further in the paper, the Unidirectional LSTMs and GRUs are extended to their bidirectional counterparts, which train input data in both the forward and reverse directions. The work presented in this paper clearly indicates a superior performance of the bidirectional models compared to their unidirectional counterparts. The bidirectional LSTM shows an accuracy of about 90% with a nominal loss of 0.0388 for the same training iterations and learning rate and being tested on a well established english speech commands data-set.

Read the paper · More papers on PaperTik