High Performance and Energy Efficient Integer Matrix Multiplication for Deep Learning

Pau San Juan, Pedro Alonso, Enrique S. Quintana–Ort́ı · 2021

We present a multi-threaded implementation of the matrix multiplication for deep learning on ARM multicore processors. Following standard practice for inference with convolutional neural networks, our GEMM kernel operates with 16-bit integer arithmetic, yielding significant performance acceleration and cutting the memory requirements with respect to IEEE (floating point) single precision by half, allowing the deployment of larger neural network models on low power devices with limited storage capacity.

Read the paper · More papers on PaperTik