A Scalable Hardware Architecture for Efficient Learning of Recurrent Neural Networks at the Edge

Yicheng Zhang, Manil Dev Gomony, Henk Corporaal, Federico Corradi · 2024

Edge devices can execute pre-trained Artificial Intelligence (AI) models optimized on large Graphical Processing Units (GPU) but often need fine-tuning for real-world data. This process, known as edge learning, is crucial for personalized learning for tasks such as speech and gesture recognition and often requires recurrent neural networks (RNNs). However, training RNNs on edge devices faces challenges due to limited resources. We propose a system for RNN training through sequence partitioning using the Forward Propagation Through Time (FPTT) training method, facilitating edge learning. Our optimized HW/SW co-design for FPTT is the first of its kind. In our work, we have implemented the complete computational process for training Long Short-Term Memory (LSTM) networks using FPTT, and we have optimized and explored the hardware architecture leveraging the Chipyard framework. Our findings indicate considerable memory savings, with only a slight increase in latency, when training small-batch size sequential MNIST (S-MNIST) data.

Read the paper · More papers on PaperTik