FullPack: Full Vector Utilization for Sub-Byte Quantized Matrix-Vector Multiplication on General Purpose CPUs
Hossein Katebi, Navidreza Asadi, Maziar Goudarzi · IEEE Computer Architecture Letters · 2024
Sub-byte quantization on popular vector ISAs suffers from heavy waste of vector as well as memory bandwidth. The latest methods pack a number of quantized data in one vector, but have to pad them with empty bits to avoid overflow to neighbours. We remove even these empty bits and provide full utilization of the vector and memory bandwidth by our data-layout/compute co-design scheme. We implemented FullPack on TFLite for Vector-Matrix multiplication and showed up to$6.7\times$speedup,$2.75\times$on average on single layers, which translated to$1.56-2.11\times$end-to-end speedup on DeepSpeech.