Temporal Staggering of Applications Based on Job Classification and I/O Burst Prediction
Wenxiang Yang, Cheng Chen, Jie Yu · 2022
Large-scale or ultra-scale applications in the supercomputers typically spend a significant fraction of the execution time performing I/O operations, however, I/O traffic varies on different time slices. These peaks in I/O traffic that pop up in the time dimension are usually called the I/O burst. I/O bursts from multiple jobs flooding the system together may cause severe I/O contention, and further degrade the performance of the system and application. Therefore, if these concurrent I/O bursts can be staggered from each other, the possibility of conflict will be greatly reduced. Unfortunately, although I/O scheduling tends to be an effective way for approaching the inter-job I/O interference issue, it has never been available on production supercomputers, even there is rare discussion on academic study. We collect historical job logs and node-side I/O traces, which are jointly to identify the I/O patterns of each job. Typically, similar I/O behaviors are observed across various jobs with similar features, especially, the write I/O bursts in a job exhibits repetitive patterns and is predictable. We predict the I/O bursts of the job and propose a staggered scheduling policy to reduce the possibility of multiple bursts stepping on each other's toes. Extensive evaluations are performed on a production supercomputer and the burst-aware staggered scheduling can indeed relieve the I/O conflict and save I/O time of the application significantly.