Astro: Auto-Generation of Synthetic Traces Using Scaling Pattern Recognition for MPI Workloads
Jian Chen, Russell M. Clapp · IEEE Transactions on Parallel and Distributed Systems · 2017
Performance modeling of scale-out MPI workloads is critical in assessing trade-offs for high-performance system designs. Traces and workload skeletons are the two main vehicles used to date to accomplish this task. However, the ever-increasing scale and complexity of MPI workloads makes it difficult, sometimes infeasible, to collect the traces and/or skeletonize the workloads, due to the constraints of computing resources or the unavailability of workload source code. This paper presents Astro, a framework that leverages machine learning techniques to automatically recognize the scaling patterns from training traces, and generate high-quality synthetic traces which mimic original trace behavior and extrapolate it to arbitrary scale. Experimental results show that compared with original traces, the synthetic traces yield less than 15 percent error against a range of metrics for up to 8 K MPI ranks. This framework enables large-scale performance modeling with limited computing resources, and allows modeling proprietary workloads in a portable and secure way.