dMath: Linear Algebra for Scaleout GP-GPUs

Steven Eliuk, Cameron Upright, Anthony Skjellum · 2016

A new scalable parallel math library, dMath, is presented that demonstrates leading scaling when using intranode, internode, and hybrid-parallelism for deep learning (DL). dMath provides easy-to-use distributed primitives and a variety of domain-specific algorithms. These include matrix multiplication, convolutions, and others allowing for rapid development of scalable applications, including Deep Neural Networks (DNNs), whereas previously one was restricted to libraries that provided effective primitives for only a single GPU, like Nvidia's cuBLAS and cuDNN or DNN primitives from Nervana's neon framework. dMath enables a wide range of developers to utilize parallel and distributed hardware easily. Data is stored persistently on the GPU hardware, avoiding costly transfers between host and device. Advanced memory management techniques are utilized, including caching of transferred data and memory reuse through pooling. dMath delivers performance, portability, and productivity to its specific domain of support. dMath's caching approach addresses one of the key drawbacks of GPUs, which is to keep data sets cached and to avoid overheads of the CPU-GPU memory interface wherever possible.

Read the paper · More papers on PaperTik