Benchmarking performance of RaySGD and Horovod for big data applications
Shruti Kunde, Amey Pandit, Rekha Singhal · 2020
With the advent of big data, training deep learning models quickly has gained prime significance. The faster a model is trained, the more relevant are its predictions in a given context. Deep learning is used for non structured data such as images, videos, sounds, text corpus, all of which represent a huge volume of data and also use complex models. Training these workloads can often take days or even weeks, because of various factors such as size of data, complexity of model, network and the underlying hardware infrastructure. The recognized divide and conquer solution to expedite the training process is to distribute either data or the model. Alas, the challenges of a distributed training setup are well known - creating and maintaining a cluster, enabling data or model parallelism along with uninterrupted communication across the cluster nodes.In this paper, we focus on two lightweight libraries for distributed deep learning, RaySGD and Horovod, which aim to alleviate these challenges by providing support for seamless parallellization. We conduct an in-depth benchmarking exercise to evaluate the performance of both libraries for training time(latency) incurred. Our experiments are conducted on a combination of various parameters such as hardware setup (CPU or GPU based), standard and manually coded models, real world and synthetic datasets. We also vary batch sizes of large workloads and number of worker nodes in a distributed setting. The insights obtained from our experiments act as guidelines for data scientists, facilitating the decision making process when conducting distributed training of big data applications on RaySGD or Horovod.