Self-organizing infrastructure for machine (deep) learning at scale
Samir Mittal · 2018
Building machine (deep) learning1 systems is hard. Computation requirements grow non-linearly with the complexity of the task at hand creating acute challenges relating to data dimensionality, complex model development, slow experiments, and scalability of production deployments. The bulk of the ML/DL effort is consumed in infrastructure and data management. Automating such workflows has become the focus of recent research activity, so as to make ML/DL systems universally accessible. We extend these paradigms by infusing domain knowledge for infrastructure self-management. Key elements include understanding application design intent, fingerprinting the neural network for its computational, data and convergence properties, optimizing the implementation to achieving workload intent, and accelerating the neural network implementation in real-time hardware implementation. Keys to success require offline behavioural modelling coupled with online dynamic adaptation, made possible by the use of cognitive algorithms that accumulate knowledge in a dynamic and continuously evolving knowledgebase. In this way, we use machine learning to automate AI infrastructure management to minimize human engineering, and significantly accelerating application performance.