Choosing Training and Serving Infrastructure
Dan Sullivan · 2020
This chapter focuses on choosing the appropriate training and serving infrastructure for your needs when serverless or specialized AI services are not a good fit for your requirements. Advances in integrated circuit and instruction set design have led to the development of specialized computing hardware accelerators. These devices offload some of the computing workload from CPUs. Graphic processing units (GPUs) are accelerators that have multiple arithmetic logic units (ALUs), which implement adders and multipliers. Modern GPUs have thousands of ALUs. This architecture is well suited to workloads that benefit from massive parallelization, such as training deep learning models. GPUs typically increase training performance on deep learning models by a factor of 10. Tensor processing units (TPUs) are specialized accelerators based on ASICs and created by Google to improve the training of deep neural networks. These accelerators are designed for the TensorFlow framework. The design of TPUs is proprietary and available only through Google's cloud services or on Google's Edge TPU devices. Single machines are useful for training small models, developing machine learning applications, and exploring data using Jupyter Notebooks or related tools. Cloud Datalab, for example, runs instances in Compute Engine virtual machines (VMs).