Quantifying Performance Gains of GPUDirect Storage
Devasena Inupakutika, Bridget Davis, Qirui Yang, Daniel Kim, David Akopian · 2022
Rapid growth in data collection has led to a need for reconditioning the underlying computation and storage solutions. In data-intensive workloads involving machine learning and large-scale simulations, the computations are shifting from CPUs to GPUs. The overhead of input and output (IO) operations in these workloads during the data transfer between GPU and storage becomes a new performance bottleneck. It has thus become evident that the traditional approach of data transfers between GPU memory and device storage that involves CPU as a buffer, has limited GPUs' ability to utilize their vast resources efficiently. Additionally, the bounce buffer approach wastes CPU cycles spent on transferring data. These paths introduce latency, and reduce overall GPU processing performance, especially if the CPU does not need to utilize the data. To solve the resulting bottlenecks, a new technology, NVIDIA GPUDirect Storage (GDS) supported on NVIDIA GPU, accelerates GPU-storage communication by establishing a direct path between local NVMe or remote storage and GPU memory. In this work, we quantify high throughput gains, low latency achievements, and CPU utilization savings of GDS technology with Weka as remote storage cluster. We experiment with various workloads and identify workloads that benefit the most from these employed technologies. We systematically measure the performance of synthetic representative data read and real machine learning workloads optimized with GDS. Finally, we utilize the findings to discuss correlated implications to systems with GDS. Demonstration results are twofold: a) A testing methodology for the performance evaluation of a GPU client with GDS supported file systems for local NVMe and Weka, b) Establish a baseline level of GDS performance for local and remote storage devices.