Orchestrating deep learning workloads on distributed infrastructure
Seetharami Seelam, Yubo Li · 2017
Containers simplify the packaging, deployment and orchestration of diverse workloads on distributed infrastructure. Containers are primarily used for web applications, databases, application servers, etc. on infrastructure that consists of CPUs, Memory, Network and Storage. Accelerator hardware such GPUs are needed for emerging class of deep learning workloads with unique set of requirements that are not addressed by current container orchestration systems like Mesos, Kuberentes, and Docker swarm. In this extended abstract, we discuss the requirements to support GPUs in container management systems and describe our solutions in Kubernetes. We will conclude with a set of open issues that are yet to be addressed to fully support deep learning workloads on distributed infrastructure.