MemCore: Computing-in-Flash Design for Deep Neural Network Acceleration

Shaodi Wang · 2022 6th IEEE Electron Devices Technology & Manufacturing Conference (EDTM) · 2022

Neural Networks (NNs) have been widely employed in modern artificial intelligence (AI) systems due to their unprecedented capability in classification, recognition and detection. However, the massive data communication between memory and processing units has been proven to be the main challenge to improve the efficiency of NNs based hardware. Significant power and latency are demand for NNs inference. Advanced process and package can mitigate the memory wall problem at the expense of cost. Computing-in-memory (CIM) is potentially the most efficient solution to this problem. However, CIM is an non Vonneuman architecture and uses analog signal instead of logical one. Classic methodologies of design, process, reliability, precision and compiler are impractical. We have explored a CIM architecture using floating technology for computation, called MemCore. MemCore supports computing precision of over 8-bits, flexible mapping of multi-layer neural networks on single core, as well as multi-core architecture to expand computing capabilities. We have released the product WTM2101 and are verifying a new product for performance over 10 Tops. This paper is a overview introduction of MemCore architecture.

Read the paper · More papers on PaperTik