Towards Scalable Deep Learning via I/O Analysis and Optimization

Sarunya Pumma, Min Si, Wu-chun Feng, Pavan Balaji · 2017

Deep learning systems have been growing in prominence as a way to automatically characterize objects, trends, and anomalies. Researchers have been investigating techniques to optimize such systems. An area of particular interest has been using supercomputing systems to quickly generate effective deep learning networks, a phase referred to as “training” of the deep neural network. As we scale deep learning frameworks-such as Caffe-on large-scale systems, we notice that parallelism can help improve the computation tremendously, leaving data I/O as the major bottleneck limiting the overall system scalability. In this paper, we present a detailed analysis of the performance bottlenecks of Caffe on large supercomputing systems. The analysis shows that Caffe's I/O subsystem-LMDB-relies on memory-mapped I/O, which can be highly inefficient on large-scale systems because of its interaction with the process-scheduling system and the network-based parallel filesystem. Based on this analysis, we present LMDBIO, an optimized I/O plugin for Caffe that takes into account the data access pattern in order to vastly improve I/O performance. Experimental results show that LMDBIO can improve the overall execution time of Caffe by nearly 20-fold in some cases.

Read the paper · More papers on PaperTik