Chameleon: Online Clustering of MPI Program Traces

Amir Bahmani, Frank Mueller · 2018

The data explosion in scientific computing applications continues to fuel increasing demand for computational power. Understanding application behavior in this context becomes essential to determine shortcomings, e.g., by collecting detailed information with tracing toolsets. This work considers parallel applications using the SPMD (single program multiple data) paradigm that relies on iterative kernels. This characteristic provides an opportunity to empower tracing toolsets with effective machine learning algorithms. One solution is to cluster processes with the same behavior into a group. Instead of collecting performance information from each individual process, this information can be collected from just a set of lead processes, i.e., one lead process per group. This work, called Chameleon, contributes an online, fast, and scalable signature-based clustering algorithm. Unlike related work, namely ScalaTrace V2, that generates the compressed global trace within MPI_Finalize, Chameleon creates this trace incrementally during the execution of applications and only for lead processes. Chameleon also identifies different program phases, clusters processes exhibiting different execution behavior, and creates a compressed global trace file on-the-fly, all incrementally at interim execution points of applications. The resulting system combines low overhead at the clustering level a lower time complexity of log (P) than prior work.

Read the paper · More papers on PaperTik