LAMBDA: An Open Framework for Deep Neural Network Accelerators Simulation
Enrico Russo, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Giuseppe Ascia, Vincenzo Catania · 2021
Many tasks in the realm of recognition, mining, and synthesis are increasingly being implemented by using machine learning approaches. In particular, deep neural networks (DNNs) are currently able to solve a myriad of tasks with superhuman accuracy. The ever more increasing cost of communication as compared to computation is pushing toward the shifting of the intelligence as close as the edge of the network. However, the edge nodes often come in the form of resource-constrained devices that do not have the computational and memory capabilities required to perform DNN inferences. To solve this problem, the use of DNN accelerators embedded into a device makes it possible to run complex DNN inferences even on resource-constrained devices. A DNN accelerator exposes many configurable hardware settings that the designer has to properly set to select the optimal tradeoff among different conflicting criteria, including performance, energy, and accuracy. This paper presents LAMBDA, a framework based on Timeloop/Accelergy infrastructure that allows exploring the design space of configurable DNN accelerators taking into account a variety of architectural and microarchitectural parameters. We describe the key components of LAMBDA focusing on the modeling of the communication and memory sub-system, which have the greatest impact on the energy and performance figures of a hardware accelerator.