NICE: A Nonintrusive In-Storage-Computing Framework for Embedded Applications

Tianyu Wang, Yongbiao Zhu, Shaoqi Li, Jin Xue, Chenlin Ma, Yi Wang, Zhaoyan Shen, Zili Shao · IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems · 2024

Embedded machine learning applications face challenges related to massive data movement and high computational intensity, exacerbated by the limited performance of mobile devices. Computational storage devices (CSDs) pose huge potential for accelerating both data-intensive and computation-intensive embedded machine learning tasks by effectively reducing data movement and leveraging built-in accelerators. However, existing in-storage-computing (ISC) frameworks either require invasive customization of existing host driver layers or necessitate complex device firmware modifications, hindering the widespread deployment of CSDs. In addition, the lack of file semantics and the constrained internal resources within CSD implicitly compromise system performance and impact normal read/write performance. In this article, we aim to provide a nonintrusive in-storage-computing framework for embedded applications, named NICE. This framework includes an easy-to-use ISC programming interface that bypasses the kernel stack and requires no modification to the host NVMe driver, which is achieved through a novel hyper-addressing-based programming library and a file-aware page data layout within the CSD. In addition, we incorporate a lightweight kernel with coroutine-based command scheduling and several FPGA-based accelerators within the storage device firmware to enhance the performance of embedded machine learning applications while ensuring that the normal I/O performance remains unaffected. NICE is implemented on real CSD hardware integrated with ARM and FPGA. Experimental results demonstrate that our NICE framework can achieve an average latency performance improvement of$43.5\times $($9.32\times $) compared to CPU-(GPU-) based embedded machine learning solutions using the state-of-the-art NVIDIA Jetson NX platform, with$27.5\times $($4.3\times $) higher energy efficiency. NICE also has$34.2\times $less software and I/O performance overheads than state-of-the-art ISC frameworks.

Read the paper · More papers on PaperTik