Lotus: Characterization of Machine Learning Preprocessing Pipelines via Framework and Hardware Profiling

Rajveer Bachkaniwala, Harshith Lanka, Kexin Rong, Ada Gavrilovska · 2024

Preprocessing input data is a crucial step in machine learning pipelines, involving tasks such as loading, decoding, and applying transformations. Prior works have identified preprocessing as a performance bottleneck for ML training jobs and introduced optimizations like CPU and I/O parallelism, accelerator offloading, and on-node/distributed computing as well as caching to mitigate this bottleneck. However, there is a lack of support for characterizing preprocessing pipelines at a finer granularity, especially at the microarchitecture level, which can provide insights to validate and inform the design of existing and future optimization techniques.To enable these insights, we introduce Lotus, a profiling tool for the preprocessing stage of ML pipelines. Firstly, it captures fine-grained preprocessing events (e.g., <10 ms) with minimal time and storage overheads. Secondly, it bridges the gap between high-level Python functions and low-level hardware performance counters by reconstructing a mapping between Python functions and the underlying C++ functions they invoke. This unique combination enables users to better reason about their pipeline’s performance at both the framework and CPU architecture levels. We demonstrate the insights made possible by applying Lotus to representative ML workloads and compare its capabilities, overheads, and ease of use with alternative profilers.

Read the paper · More papers on PaperTik