INSPIRE: Accelerating Deep Neural Networks via Hardware-friendly Index-Pair Encoding
Fangxin Liu, Ning Yang, Zhiyan Song, Zongwu Wang, Haoming Li, Shiyuan Huang, Zhuoran Song, Songwen Pei, Li Jiang · 2024
Deep Neural Network (DNN) inference consumes significant computing resources and development efforts due to the growing model size. Quantization is a promising technique to reduce the computation and memory cost of DNNs. Most existing quantization methods rely on fixed-point integers or floating-point types, which require more bits to maintain model accuracy. In contrast, variable-length quantization, which combines high precision for values with significant magnitudes (i.e., outliers) and low precision for normal values, offers algorithmic advantages but introduces significant hardware overhead due to variable-length encoding and decoding. Also, existing quantization methods are less effective for both (dynamic) activations and (static) weights due to the presence of outliers.