Memory Efficient Training using Lookup-Table-based Quantization for Neural Network

Kazuki Onishi, Jaehoon Yu, Masanori Hashimoto · 2020

Modern neural networks require a tremendous number of parameters, which causes unaffordable requirements of memory and computation resources for embedded systems. To tackle this issue, we propose a LUT-based training method in this paper. The proposed method consists of two components: cluster swap and factorization. Cluster swap is an extension to the quantization process in Deep Compression that overcomes its drawback of unstable training by reassigning each parameter to the proximate cluster. Factorization is a computation trick to reduce the computational cost of neural networks. The experimental results show that the proposed method can decrease memory usage for forward and backward processes to 22.2% and 60.0%,respectively, and reduce the number of multiplications to 11.7%, with 1.41% accuracy loss.

Read the paper · More papers on PaperTik