Designing a Kubernetes Operator for Machine Learning Applications

Ali A. Kanso, Edi Palencia, Kinshuman Patra, Jiaxin Shan, Mengyuan Chao, Wei Zhi Xu, Tengwei Cai, Kang Chen, Shuai Qiao · 2021

Machine Learning workloads such as deep learning and hyperparameter tuning are compute-intensive by nature. Parallel execution is key to reducing the learning time. The Ray Framework is a distributed middleware that provides primitives to seamlessly parallelize machine learning code execution across a cluster of compute node. Launching a Ray managed machine learning application requires a Ray cluster that is diligently configured, well connected and easily scalable. Kubernetes, the container management middleware, satisfies all the requirements to create and scale ray clusters. However, setting up a cluster within Kubernetes, is a tedious and error prone task when done manually. In this paper we present KubeRay, an Operator and suite of tools designed, and built to create Ray cluster in Kubernetes with minimum effort. We present our architectural choices, our open-source implementation, and we analyze the performance of our solution.

Read the paper · More papers on PaperTik