Accelerating Deep Learning Workloads with Advanced Matrix Extensions

Wenhuan Huang, Duyi Wang, Shan Zhou, Changqing Li, Yi Xie, Pujiang He · 2023

This paper introduces advanced matrix extensions which is a new x86 extension that is designed to accelerate deep learning performance. It discusses the evolution of hardware and the theoretical peak performance of the accelerator. Two approaches for utilizing this technology to expedite deep learning workloads are suggested: one involves re-implementing the code, while the other entails adopting optimized libraries. The experimental findings demonstrate a substantial 1.99x performance enhancement of the Wide & Deep model on the Sapphire Rapids platform by using advanced matrix extensions.

Read the paper · More papers on PaperTik