Accelerating the Deep Learning Recommendation System Model Inference on X86 Processors

Wenhuan Huang, Changqing Li, Shan Zhou, Pujiang He · 2023

Deep learning recommendation systems have commonly been regarded as an indispensable part of E-commerce. The exponential growth of model size requires a growing demand for hardware. In this paper, the overview of the recommendation system and three significant hardware features - Advanced Vector Extensions (AVX), Brain Floating Point (BFloat16), and Advanced Matrix Extensions (AMX) were respectively introduced. Then this paper introduces multiple effective and generic optimization methods for accelerating deep learning recommendation system models' inference performance on Intel® Xeon Scalable processors which take full advantage of the capabilities of the hardware. Developers could leverage these methods on their own recommendation systems and get significant improvements in performance and productivity.

Read the paper · More papers on PaperTik