Using Luong and Bahdanau Attention Mechanism on the Long Short-Term Memory Networks
Vinayak Ashok Bharadi, Sujata Alegavi · 2021
Recurrent neural networks (RNNs) are popular for processing data of sequential nature and sequence modelling problems. Long Short-Term Memory (LSTM) networks are a special type of RNNs, they have addressed the problems faced by RNNS and have been proven successful in sequence prediction and analysis. The Encoder-Decoder LSTMs are designed to address sequence-to-sequence problems, also referred to as seq2seq. As the length of the input and output sequence increases, it results in the degradation of the performance of the seq2seq LSTM model. The attention mechanism enables a Deep Neural Network to selectively concentrate on a few relevant things. It enables the decoder to cater to the different parts of the source sequence for each time step of the output. This adds context for each time step. The Luong attention mechanism employs the top hidden layer states of both the encoder and decoder, whereas the chain of forward and backward hidden states of the source is taken by the Bahdanau attention mechanism. In this work, the Luong and Bahdanau attention mechanisms are implemented on Encoder-Decoder LSTMs and their performance is evaluated. The attention mechanism shows better performance than the regular network while predicting the COVID-19 deaths.