Generative AI-Powered Spark Cluster Recommendation Engine

Arpana Dipak Mahajan, Akshay Mahale, Amol S. Deshmukh, Arun Vidyadharan, Vijeth S Hegde, Koushik Vijayaraghavan · 2023

Apache Spark is a data processing framework that performs complex processing tasks on very large data sets very smoothly. It distributes data processing tasks across multiple compute instances, either on its own or in tandem with other distributed computing tools. With the increasing amount of data and evolution of ML models, the need to accomplish complex feature engineering/pre-processing, Model training quickly has become optimum. As compared to a single compute instance, a cluster of compute instances shows a remarkable performance boost to achieve faster data processing. Since a cluster is a combination of multiple compute instances (Worker Nodes) managed by a Master Node, the total cost to utilize such cluster configuration is very high. But it is seen that there is significant cost reduction as these clusters like (Databricks, and EMR) are governed by the Cloud platform, there are the options of paying as you use. Also depending on workload, the cluster can be designed in such a way that it provides the best performance at lower costs. But this process is manual and needs sufficient technical skills and experience to design a cluster. In the presented approach, automation of the cluster selection process is demonstrated, by developing a GEN-AI-based recommendation engine. This work presents Conditional generative adversarial networks (cGANs) which are used to produce more samples from the joint distribution of sparse custom training data. Depending upon the data workload, expected usage time, and budget presented recommendation engine suggests instance type for master and worker nodes along with the number of worker nodes needed.

Read the paper · More papers on PaperTik