Big Data Acceleration on Tightly Coupled FPGA: An Alternating Least Squares inference Case Study

Piotr Ratuszniak, Srivatsan Krishnan, Ilya Burylov, Ramesh Illikkal, Chong Yeol Nah, M. Łącki · 2025

The workloads that run on the cloud are getting increasingly diverse and require customization for better performance and energy efficiency. On the other hand, general purpose and high performance Xeon®processors are ubiquitous in cloud and data centers. Barriers to Dennards scaling and emergence of the dark silicon issues are forcing architects to look into adding application-specific accelerators to squeeze as much energy efficiency and performance as possible. In such a scenario FPGAs provide an optimal middle ground not only with performance and energy efficiency but also as a highly customizable hardware accelerator. As Intel®couples high performance Xeon®processors with highly customizable and energy efficient FPGA, it provides unique challenges and opportunities for mapping cloud scale workloads on these heterogeneous architectures. In this paper, the challenges of enabling a popular big data framework called Apache Spark for the first generation Intel®Xeon®with integrated FPGA are presented. The FPGA acceleration capability of a big data application on the integrated platform through Intel®Data Analytics and Acceleration Libraries are exposed. A novel load balancing scheme for integrated FPGA is presented that allows programmers to view an FPGA accelerator as general purpose compute resource. In doing so, the complexity of programming the FPGA is completely abstracted from the programmer. A fully functional end-to-end inference application with FPGA acceleration, measuring speedup of $1.5 \mathrm{x} /$ socket as seen from Apache Spark when compared to the optimized Xeon®implementation is shown.

Read the paper · More papers on PaperTik