AIO: Automating I/O Optimization Pipeline for Data-Intensive Applications in HPC
Wanxin Wang, Huijun Wu, Lihua Yang, Zhangyu Liu, Bin Zheng · 2024
High-performance computing (HPC) systems has entered the exascale era, but I/O performance has lagged behind due to storage hardware limitations, creating a "storage wall effect" that hinders HPC systems full potential. Modern HPC storage systems employ a layered storage architecture with local node storage, burst buffers, and global parallel file systems to improve I/O performance. However, application workloads vary, and fast storage layers may not always outperform global parallel file systems, while a static storage software stack configuration may not achieve optimal performance across different applications. Traditional I/O optimization is time-consuming, laborintensive, and prone to human error. To address this issue, this paper proposes AIO, a data-driven I/O optimization method. AIO automatically selects the optimal storage hierarchy and software stack based on the task's I/O patterns, and further enhances performance by identifying and applying the most efficient configuration settings. We evaluated AIO with eight applications in a cluster that features shared burst buffers, local burst buffers, and a global shared file system. The experiments demonstrate that AIO effectively optimizes I/O performance, achieving an overall speedup of 2.74.