A PMIC-Based Component-Level Power Profiling Framework for On-Device AI Optimization
Won-Seok Chang, Seung-Ryeol Ohk, Yong Wook Kim · IEEE Access · 2026
While model compression is crucial for on-device AI, conventional proxy metrics such as FLOPs and latency often fail to reflect actual energy consumption owing to complex system-on-chip (SoC) interactions involving memory hierarchies, interconnects, and dynamic voltage and frequency scaling (DVFS). Moreover, existing measurement techniques lack the temporal resolution and component-level decomposition required to analyze layer- and stage-wise energy variations. This paper proposes a noninvasive profiling framework that decomposes the power consumption across the CPU, GPU, TPU, and DRAM without external equipment. By combining power-rail-to-component mapping, PMIC-based power acquisition, and statistical alignment with time-weighted attribution, the framework constructs stage- or layer-level energy profiles, even under low sampling rate conditions, adapting to delegate-specific execution characteristics. The evaluation of quantized and pruned ResNet50 variants revealed that the energy impact of identical compression techniques varied substantially by delegate and precision. Under CPU FP32, early-stage pruning yielded the largest savings, whereas stem pruning incurred a penalty; under CPU INT8, only late-stage pruning offered marginal benefits. GPU results showed mixed outcomes under FP32 but overall reductions under INT8, whereas TPU efficiency varied sharply depending on the pruning location, with late-stage and global-pruning proving beneficial, whereas stem- and middle-stage pruning increased consumption. These findings underscore the fact that no single proxy metric reliably predicts energy efficiency, necessitating component-level measurements using real hardware. This study provides a quantitative analysis and data-driven guidelines for effective on-device energy optimization.