Fast convolutions leveraging pruning aware runtime code generation
Malith Jayaweera Don Satarasinghege · 2022
Accelerating convolutions have been pursued by the research community, especially given their wide-spread use in Deep Neural Networks (DNNs) in Machine Learning (ML) applications. Toom Cook's algorithm, also widely referred to as Winograd convolution, has been a centerpiece in speeding up 3X3 convolutions on the CPU, as well as on the GPU. Although, Winograd convolutions have been optimized further through just-in-time (JIT) code generation, current approaches have been restricted to only instruction sequencing optimizations. In our work, we combine the application of pruning of DNNs in a structured manner with JIT-based code generation, automatically detecting sparsity patterns and generating highly optimized instruction sequences to further improve the performance of Winograd convolutions. When working with layers that exhibit 60% and 80% channel sparsity, we achieve a geometric mean speedup of 1.26X and 1.52X, respectively, for 3X3 kernel convolutions. We intend to open source our code implementation, enabling the broader community to leverage our code generation techniques and produce high performance convolutions on CPU platforms.--Author's abstract