G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor Migrations
H. Zhang, Yirui Zhou, Yuqi Xue, Yiqi Liu, Jian Huang · 2023
To break the GPU memory wall for scaling deep learning workloads, a variety of architecture and system techniques have been proposed recently. Their typical approaches include memory extension with flash memory and direct storage access. However, these techniques still suffer from suboptimal performance and introduce complexity to the GPU memory management, making them hard to meet the scalability requirement of deep learning workloads today.