Dive Into Recurrent Neural Networks
Mohamed Abdel‐Basset, Nour Moustafa, Hossam Hawash · 2022
The notable limitation of recurrent networks resides in their numerical instability. Even though gradient clipping could provide a tricky solution to alleviate this issue, the vanilla recurrent neural network (RNN) is unable to retain the sequential information in its memory state in case of the too long input sequence. This chapter provides a deep dive into such widely used recurrent networks, starting with long short-term memory (LSTM). It illustrates the structure of the LSTM cell supplemented with “peephole connections”. The chapter explores the gated recurrent units (GRUs) in terms of design principles, working methodology, and steps of the calculation, and also describes the main points of similarities and differences with LSTM and vanilla RNN. It provides a basic understanding of deep recurrent networks that stack multiple recurrent hidden layers. The main topics covered (e.g. LSTM, GRU, convolutional LSTM, etc.) are practically validated and implemented for detecting malware in the real-world IoT.