FHEDGE: Encrypted Inference on Lightweight Edge Devices
Soumik Sinha, Sayandeep Saha, Ayantika Chatterjee, Debdeep Mukhopadhyay · 2024
Deployment of Fully Homomorphic Encryption (FHE) enabled Machine Learning (ML) inference has been largely restricted to high end public cloud servers due to their formidable computational overheads. In this work, we propose FHEDGE, a first-of-its-kind encrypted inference engine which achieves an end-to-end image classification task within a few hours on a cluster of low-power embedded devices. The major computational bottleneck of FHE workloads is repeated invocations of bootstrapping operations, which are used for denoising of ciphertexts. To address this, we propose an efficient neural network weight encoding approach which significantly reduces bootstrapping overhead of encrypted multiplications required to realize the Multiply-and-Accumulate (MAC) units of inference architecture. We further show how computational and memory overhead of residual MAC components can be optimized by mapping them into Single-Instruction-Multiple-Data (SIMD) style execution templates, which are inherently data parallel abstractions of homomorphic circuits. When evaluated on a cluster of 4 low end Raspberry Pi devices, our proposed techniques make it feasible to run an end-to-end MNIST classification task within 8.5 hours, thus yielding speedup of $7.2 \times$ and higher throughput of $1.77 \times$ over the considered baselines. Memory footprints ($1.7 G B$) of FHEDGE shows its suitability towards resource constrained analytics use-cases.