Quantization and Pruning of Multilayer Perceptrons: Towards Compact Neural Networks
Tomas Lundin, Perry D. Moerland · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 1997
A connectionist system or neural network is a massively parallel network of weighted interconnections, which connect one or more layers of non-linear processing elements (neurons).To fully prot from the inherent parallel processing of these networks, development of parallel hardware implementations is essential.However, these hardware implementations often dier in various ways from the ideal mathematical description of a neural network model.It is, for example, required to have quantized network parameters, in both electronic and optical implementations of neural networks.This can be because device operation is quantized or a coarse quantization of network parameters is benecial for designing compact networks.Most of the standard algorithms for training neural networks are not suitable for quantized networks because they are based on gradient descent and require a high accuracy of the network parameters Several weight discretization techniques have b e e n d e v eloped to reduce the required accuracy further without deterioration of network performance.One of the earliest of these techniques Fiesler-88] is further investigated and improved in this report.Another way to obtain compact networks is by minimizing their topology for the problem at hand.However, it is impossible to know a priori the size of such minimal network topology.An unsuitable topology will increase the training time, lower the generalization performance on unseen test data Gosh-94], and in some cases even cause non-convergence.One method to lower the importance of choosing the initial network topology and minimizing the network size is pruning, that is, removal of connections or neurons during training.Especially a combination of parameter/weight q u a n tization and network pruning, leading to networks that have a small topology and for which small accuracy is su cient, is of great importance for hardware implementation of neural networks.Such n e t works oer a minimization of chip area and computational requirements.Due to their lack of redundancy they are also expected to show a better generalization on unseen patterns (Occam's razor).Such a combination of pruning techniques with weight q u a n tization is studied in the second part of this report.Five dierent quantization functions, chapter 3, and six pruning methods, chapter 4, are evaluated in a series of experiments on real-world benchmarks problems.The main goal is to rst improve the original weight discretization technique as much as possible to obtain networks with both a small number of discrete weight levels and good generalization performance on unseen test data.Secondly, the results from the rst part are combined with the pruning methods to ease the choice of the initial network topology and obtain compact networks.