LEAP: Lightweight Neural Network Inference Through Proactive Early-Exiting Prediction
Yingtao Shen, Xiangjie Li, Yehan Ma, Weidong Cao, Jie Zhao, An Zou · IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems · 2025
In recent years, the incorporation of early exit layers into deep neural networks has allowed inference to terminate earlier while maintaining accuracy. However, the passive decision-making involved in the these static exit placement creates a dilemma: fine-grained placement may cause high performance and energy overhead due to frequent exit layer execution, while coarse-grained placement may miss early exit opportunities. Moreover, common energy-saving techniques like adjusting processor configurations are not applicable once inference begins. To overcome these challenges and improve computation and energy efficiency, we propose LEAP, a software-hardware co-design approach. On the software side, LEAP proactively predicts exit points at runtime, reducing computation by enabling early exits without requiring every pre-placed exit layer to be executed. On the hardware side, LEAP adjusts processor settings—such as frequency and voltage—based on single or multiple predicted exits to optimize energy consumption while adhering to latency requirements. Extensive experimental results show that LEAP significantly improves efficiency. Compared to standard inference, LEAP reduces computation by up to 76.4% and saves up to 83.2% in energy. Compared to state-of-the-art early exit methods, LEAP achieves up to 27.9% less computation and 57.1% more energy savings, while maintaining similar accuracy and latency.