STGM: Spatio-Temporal GPU Management for Real-Time Tasks
Sujan Kumar Saha, Yecheng Xiang, Hyoseung Kim · 2019
Graphics Processing Units (GPUs) have been considered as a promising technology to address the high computational demands of real-time data-intensive applications. Today's embedded processors already offer on-chip GPUs, the use of which can greatly help satisfy the timing requirements of realtime tasks by accelerating their execution. However, existing GPU management schemes either underutilize the GPU due to strictly serialized execution or introduce non-deterministic delay caused by uncontrolled concurrent execution. In this paper, we present a spatial-temporal GPU management framework that controls the allocation and sharing of GPU's internal execution engines, e.g., streaming multiprocessors in Nvidia architectures, with analytical bounds. This approach allows multiple GPU-using tasks to simultaneously execute on the GPU, thereby improving GPU utilization and reducing the worst-case response time. Also, it can improve temporal isolation by allocating a portion of GPU execution engines to tasks for their exclusive use. We have examined the feasibility of our framework on two Nvidia GPUs: GTX970 and AGX Xavier. Experimental results with randomly-generated tasksets indicate that our framework yields a significant benefit in schedulability compared to the existing real-time GPU management approaches.