Syncfree optimizers and compiler improvements for efficient model training
Shreyas Subramanian, Corey D Barrett, Yanming Wang, Ricky Das, Harish Tummalacherla, Lokeshwaran Ravi, Vinay Kumar Burugu, Yi‐Hsiang Lai · 2023
Deep learning training compilers accelerate and achieve more resource-efficient training. We present a deep learning compiler for training consisting of three main features, a syncfree optimizer, compiler caching and multi-threaded execution. We demonstrate speedups for common language and vision problems against native and XLA baselines implemented in PyTorch.