Scaling deep learning without increasing batchsize
Alexander Douglas Heye · Concurrency and Computation Practice and Experience · 2019
Summary Deep learning has proven itself to be a difficult problem in the HPC space. Although the algorithm can scale very efficiently with a sufficiently large batchsize, the efficacy of training tends to decrease as the batchsize grows. Scaling the training of a single model may be effective in narrow fields such as image classification, but more generalizable options can be achieved when considering alternate methods of parallelism and the larger workflow surrounding neural network training. Hyperparameter optimization, data set segmentation, hierarchical fine tuning, and model parallelism can all provide significant scaling capacity without increasing batchsize and can be paired with a traditional, single‐model scaling approach for a multiplicative scaling improvement. This paper intends to further define and examine these scaling techniques in how they perform individually and how combining them can provide significant improvements in overall training times.