Compute cache for data parallel acceleration

Reetu Das · 2019

The talk will start with an overview of our work Neural Cache architecture which is capable of fully executing convolutional, fully connected, pooling layers in-cache and also supports quantization in-cache. Then I will present a versatile Compute Cache architecture named Duality Cache, which re-purposes cache structures to transform them into massively parallel compute units capable of running arbitrary data parallel workloads including Deep Neural Networks. Our work presents a holistic approach to building Compute Cache system stack with techniques of performing in-cache floating-point and fixed-point arithmetic, transcendental functions, enabling SIMT execution model, designing a compiler that accepts existing CUDA programs, and providing flexibility in adapting for various workload characteristics.

Read the paper · More papers on PaperTik