Binary Recurrent Unit: Using FPGA Hardware to Accelerate Inference in Long Short-Term Memory Neural Networks
Thomas C. Mealey · OhioLink ETD Center (Ohio Library and Information Network) · 2018
Long Short-Term Memory (LSTM) is a powerful neural network algorithm that has been shown to provide state-of-the-art performance in various sequence learning tasks, including natural language processing, video classification, and speech recognition.Once an LSTM model has been trained on a dataset, the utility it provides comes from its ability to then infer information from completely new data.Due to the large complexity of LSTM models, the so-called inference stage of LSTM can require significant computing power and memory resources in order to keep up with a real-time workload.Many approaches have been taken to accelerate inference, from offloading computations to GPU or other specialized hardware, to reducing the number of computations and memory footprint required by compressing model parameters.This work takes a two-pronged approach to accelerating LSTM inference.First, a model compression scheme called binarization is identified to both reduce the storage size of model parameters and to simplify computations.This technique is applied to training LSTM for two separate sequence learning tasks, and it is shown to provide prediction performance I have greatly enjoyed my foray into the field of deep learning over the past year.First and foremost, I would like to thank my wife, Michelle, for her support and encouragement throughout the process.Without your help, this thesis would not have been possible.I would also like to thank my