TinyMem: Boosting Multi-DNN Inference on Tiny AI Accelerators with Weight Memory Virtualization
Changmin Jeon, Taesik Gong, Juheon Yi, Fahim Kawsar, Chulhong Min · 2025
As wearable devices continue to integrate deeper into our everyday lives, the importance of tiny AI accelerators in enabling efficient multi-DNN inference becomes increasingly evident. However, we identified a critical bottleneck in deploying multi-DNN models on these accelerators: weight memory loading time. This challenge is exacerbated by the unique characteristics of tiny AI accelerators, including a 2-dimensional weight memory layout and heterogeneous processors.