Accelerating K-means clustering on a tightly-coupled processor-FPGA heterogeneous system
Tarek S. Abdelrahman · 2016
We present a case study of the design of an FPGA accelerator for a tightly-coupled shared-memory processor-FPGA system: the Intel QuickAssist FPGA platform. We use K-means as an example computationally-intensive application and design a pipelined accelerator for calculating minimum distances between points and centroids. Our accelerator is unique in that it works in collaboration with CPU threads, accessing shared data in system memory. It achieves a speedup of 3.8X over a reference software implementation. Moreover, the combined use of CPU threads and the FPGA accelerator achieves a speedup of up to 2.9X over using only the CPU threads and of up to 1.9X over using only the accelerator, depending on the number of CPU threads. We analyze the impact of data sharing between the CPU threads and the FPGA and reason that while this sharing increases memory traffic, it has little impact on overall performance.