IoT-Edge Splitting With Pruned Early-Exit CNNs for Adaptive Inference
Guilherme Korol, Antonio Carlos Schneider Beck · IEEE Transactions on Very Large Scale Integration (VLSI) Systems · 2025
Deep neural networks (DNNs) are driving the Internet of Things (IoT) revolution. To manage latency and privacy concerns in this domain, IoT devices may offload partial or full DNN processing to nearby edge servers rather than relying only on the cloud. In this scenario, while field-programmable gate arrays (FPGAs) are efficient and provide flexibility for these resource-constrained IoT and edge devices, runtime optimizations like pruning and early exit can deliver further improvements. However, they need to be applied with careful design, since offloading and split computing may require constant synchronization of such dynamic DNN models. With that in mind, this article introduces a framework that automatically constructs inference platforms combining pruning and early exit with FPGA-based offloading. It addresses the latency-power-accuracy tradeoff, adapting inference to the unpredictable conditions of IoT-edge environments. Using convolutional neural networks (CNNs) as a case study, it achieves a reduction of up to$1.6\times $in latency and a$3.9\times $improvement in power efficiency, with minimal accuracy loss.