MSDF-SGD: Most-Significant Digit-First Stochastic Gradient Descent for Arbitrary-Precision Training

Changjun Song, Yongming Tang, Jiyuan Liu, Sige Bian, Danni Deng, He Li · 2023

Stochastic gradient descent has been a widely used machine learning algorithm, and interest in low-precision SGD is growing because it improves throughput and keeps efficient convergence. We propose MSDF-SGD, a novel approach allowing SGD to support arbitrary-precision training on FPGAs by employing most-significant digit-first arithmetic. MSDF-SGD is the first architecture that supports arbitrary-precision data, models, and intermediates at the same time. MSDF-SGD is evaluated via training linear classifiers on representative datasets. MSDF-SGD delivers a 1.6× speedup over state-of-the-art low-precision hardware implementations and converges up to 8.6× faster than cutting-edge implementations on CPUs. Finally, we provide a programming interface that permits building a custom arbitrary-precision training accelerator, making MSDF-SGD support more complicated, multi-layered and nonlinear models.

Read the paper · More papers on PaperTik