Accelerating Linear Algebra Kernels on a Massively Parallel Reconfigurable Architecture

A. Soorishetty, J. Zhou, Subhankar Pal, David T. Blaauw, H. Kim, Trevor Mudge, Ronald Dreslinski, Chaitali Chakrabarti · 2020

Much of the recent work on domain-specific architectures has focused on bridging the gap between performance/efficiency and programmability. We consider one such example architecture, Transformer, consisting of light-weight cores interconnected by caches and crossbars that supports run-time reconfiguration between shared and private cache mode operations. We present customized implementation of a select set of linear algebra kernels, namely, triangular matrix solver, LU decomposition, QR decomposition and matrix in-version, on Transformer. The performance of the kernel algorithms is evaluated with respect to execution time and energy efficiency. Our study shows that each kernel achieves high performance for a certain cache mode and that this cache mode can change when the matrix size changes, making a case for run-time reconfiguration.

Read the paper · More papers on PaperTik