Hardware Assisted Low Latency NPU Virtualization Using Data Prefetching Techniques

Jong-Hwan Jean, Dong–Sun Kim · 2025

With the rapid advancement of artificial intelligence and deep learning, the use of NPU (Neural Processing Unit) has expanded across various fields. While NPU enables faster processing of deep learning computations, model data's increasing size and complexity have increased demand for additional NPUs. NPU virtualization technology has been introduced to address this, allowing multiple deep learning applications to share a single NPU. This study proposes a method to reduce overhead in virtualized NPU environments by embedding a hardware scheduler on the NPU chip. This scheduler minimizes DRAM access latency during context switching by prefetching parts of the input and weight data for the deep learning model layers. Experimental results demonstrate that the hardware scheduler based NPU virtualization method significantly reduces memory cycles compared to conventional software-based NPU virtualization.

Read the paper · More papers on PaperTik