Design of an Optimized CMOS ELM Accelerator
Manoj Kumar Sharma, Umesh Chandra Lohani, Vivek Parmar, Manan Suri · 2019
In the last decade, artificial intelligence (AI) has emerged at the forefront of driving many technological innovations. A variety of algorithms have been proposed as possible alternatives to implement AI in computing systems. Extreme learning machine (ELM) has emerged as one of the most effective training algorithms for simple applications based on single layer feed-forward neural networks (SLFN) because of its unique training method. Hardware implementation of neural network algorithms is a critical requirement for deploying them in timesensitive applications. In this paper, we present a simplified AI accelerator based on CMOS technology that implements an ELM based inference engine. We present analysis of implementing such an accelerator on different technology nodes with a comparative analysis to analyze the impact of technology node scaling on performance of the proposed accelerator in terms of power and area. For the analysis, the workload used was a network of dimensions 81x18x1. We observed a remarkable benefit in speed (1.3x), area (14x) and power (7x) by scaling the design from 180 nm to 45 nm. Further, we present an analysis showing benefits of introducing emerging non-volatile memory (NVM) technologies like RRAM as the primary memory technology for the accelerator. The analysis shows that replacing the conventional CMOS with RRAM would give significant benefits in leakage (4.5x) and area (33x).