Host Pipes

Kyung-Ran Kang, Peter Yiannacouras · 2017

To date, OpenCL acceleration has been primarily focused on improving application level throughput. Many of the best-known practices in OpenCL sacrifice latency to do so. This inherent trade-off between throughput and latency stems from the limitation that in OpenCL, the primary way to transfer data between host and accelerator has been via global memory. In this work, we present a direct streaming interface from host CPU to OpenCL kernels, referred to as host pipes. This interface eliminates the latency overhead of waiting for data transfer to complete before execution can start on the head of that data. A prototype of this feature capable of saturating the full duplex PCIe Gen3x8 bandwidth was built on Intel's Arria 10 FPGA Development Kit.

Read the paper · More papers on PaperTik