FASCA: Framework for Automatic Scalable Acceleration of ML Pipeline
Mayank Mishra, Archisman Bhowmick, Rekha Singhal · 2021 IEEE International Conference on Big Data (Big Data) · 2021
Machine learning, a data-driven approach, is widely used to automate applications. It has been observed that 80% of the time is spent in pre-processing the data to make it available for building machine learning models. Data scientists generally develop and test these pipelines, primarily in python (a popular language for building models) for a small (in thousands) number of records or data points. These pipelines may incur a non-linear increase in execution time when used in production for large-sized data (in 10s of millions to billions of records or data points); these pipelines are not scalable in performance for larger data sizes. We have observed some common performance anti-patterns across many ML pipelines coded by ML practitioners, such as abuse of data frames and nested statements, especially in lambda functions - some of these are not perceivable on small-sized data. Once recognized, these patterns can be replaced by a high-performing piece of code to utilize the underlying hardware optimally. This paper presents a framework, FASCA, to automatically identify the significant performance bottlenecks in an ML/DL pipeline using static and dynamic analysis. FASCA executes the pre-processing pipeline on a fraction of actual data and builds a performance model to identify the top bottleneck components experiencing performance degradation on larger data sizes. Further, the framework generates a high-performing alternative to the bottleneck component using state-of-the-art techniques. We have evaluated and presented the results for the ML pipeline in the Retail domain, where we observed a non-linear degradation in performance with an increase in data size. FASCA recommends and changes a set of bottleneck components to accelerate up to 300% on larger data size (millions of records).