A Decision Support System for Automated Configuration of Cloud Native ML Pipelines

Aris Spyrou, Ioannis Konstantinou, Nectarios Koziris · 2024

Big Data systems like Apache Spark and Hadoop are cornerstones of large scale data processing. However, they are being utilized only on one or some steps of a larger data processing pipeline. Complex data pipelines that involve different datasets, infrastructures and programming libraries are simplified with the use of cloud-native tools like Kubernetes and KubeFlow, that model dependencies and manage the execution life cycle. Nevertheless, there are no complete solutions that are fully interoperable with Apache Spark and Kubeflow in an on-premises setup. In this work, we extend the Kubeflow Pipelines tool to support on-premises Apache Spark clusters. We experimentally evaluate our tool on an industry standard Big Data benchmark with various infrastructure and dataset configurations. We utilize the collected knowledge regarding Spark’s performance to train a Decision Tree-based ML system that can detect the optimal cluster configuration according to user constraints and predict query execution time. The system can be used by non-experts through a comprehensive GUI. We finally provide open-source implementations of both Apache Spark’s Kubeflow integration and the Decision Support System.

Read the paper · More papers on PaperTik