Optimal Declarative Orchestration of Full Lifecycle of Machine Learning Models for Cloud Native
Suyash Joshi, Basit Hasan, Ramasubramanian Brindha · 2024
The widespread use of Large Language Models (LLMs) in natural language processing tasks has become increasingly common. LLMs have the ability to produce content that is both coherent and contextually relevant. Notwithstanding, there exist notable obstacles in effectively implementing and overseeing these models within operational settings, especially in infrastructures that are cloud native. A cloud-native system for optimizing and delivering live learning modules on Kubernetes clusters is presented in this study. There are many obstacles in the way of effectively deploying inference services on the cloud. After obtaining a trained model, cloud operators need to specify hardware configurations for every model-serving container. These settings cover a wide range of topics, including CPU core counts, GPU specs, GPU RAM, possible setups for GPU sharing, and runtime parameters like batch size. When combined, these factors provide a vast and intricate configuration space. By utilizing the declarative resource management that Kubernetes provides, the suggested approach leverages declarative resource management capabilities of Kubernetes and allows users to easily train and serve LLMs on those clusters. Kubernetes clusters from a variety of cloud providers, including KIND (provisioned on bare metal), Google Kubernetes Engine (GKE), and Amazon Elastic Kubernetes Service (EKS), are specifically supported by the framework. This Kubernetes custom resource enables users to maximize the performance and cost-effectiveness of LLM deployments by utilizing Kubernetes’ scalability, flexibility, and resource allocation properties and allows to maintain the stable state with use of operators and watch pattern.