HA-BNN: Hardware-Aware Binary Neural Networks for Efficient Inference

Rui Cen, Dezheng Zhang, Yiwen Kang, Dong Wang · 2024

Binary Neural Networks (BNNs) use 1-bit weights and activations, replacing complex matrix multiplications with XNOR and bitcount operations, leading to a notable decrease in memory needs. It is worth highlighting that current BNNs commonly integrate floating-point or fixed-point in the first layer to maintain accuracy, causing a significant computational bottleneck. Meanwhile, many studies introduce additional floating-point operations to enhance accuracy, leading to increased utilization of hardware resources. In our study, firstly, we implement the improved thermometer encoding method to binarize the first layer, yielding a 0.9% enhancement in accuracy on ImageNet and a 50% reduction in computation and parameters, compared to the current state-of-the-art encoding method. Then, we don’t focus on improving accuracy. Instead, to further diminish inference latency, we modify the model structure from the perspective of improving hardware efficiency. With an equivalent number of parameters, the computation workload sees a 13.2% reduction compared to ReActNet-A, achieving 69.2% accuracy on ImageNet.

Read the paper · More papers on PaperTik