Memory-efficient Edge-based Non-Neural Face Recognition Algorithm on the Parallel Ultra-Low Power (PULP) Cluster
Mitul Sudhirkumar Nagar, Sayantan Maiti, Rahul Kumar, Hiren Kumar Mewada, Pinalkumar J. Engineer · 2023
Resource-constrained edge computing nodes observed scarcity in terms of memory and computation power. It is inevitable to use computer vision (CV) algorithms on the edge for feature extractions. However, it is challenging to implement compute-intensive and resource-hungry convolution neural network (CNN) algorithms on the edge. Hence, this work targets a non-neural Eigenfaces-based face recognition (FR) algorithm to reduce the memory footprint and latency on the edge. Furthermore, the effective implementation of computations related to artificial neural networks (ANN) and machine learning (ML) algorithms has recently seen the development of quantization as an important and very active area of research. Quantization of Eigenface-based FR algorithm’s training data from 32-bit floating-point to 8-bit fixed-point offers 93% accuracy on octa-core PULP platform with 26.31× less model size compared to state-of-the-art (SoA) deep learning (DL) SqueezeNet1.1-based FR on GAP8 platform. Quantization and parallelization were performed on Eigenface-based FR to improve the performance in terms of response time. Additionally, the quantized model was transferred from HyperRAM (L3) to cluster memory (L1) using direct memory access (DMA) via fabric control memory (L2) with double DMA transfer (D2T ), which reduces the number of execution cycles by 82×. With this technique, 120K faces can be recognized per second on the multi-core RISC-V PULP cluster at an accuracy of 93%. Compared to the non-DMA non-quantized (NDNQ) version of the Eigenfaces-based face recognition system, our implementation reduces the recognition time by 1742.85× and 917.43× for single- and octa-core, respectively. Quantization improves the recognition rate due to the reduction in execution cycles, instructions, and load operations by 82×, 25.37×, and 4658.74× respectively, compared to NDNQ versions of FR on the PULP cluster.