High performance asynchronous host-device communication through the Intel® FPGA host pipe extension for OpenCL™ applications
Michael Kinsner, Dirk Seynhaeve · 2018
The OpenCL™ standard defines a programming framework for heterogeneous compute engines, including a host API that manipulates device kernels. The global address space provides a mechanism for data to be passed between the host program and accelerator devices. Except for shared virtual memory (SVM) atomics in recent versions of the OpenCL standard, which are not available on many platforms, data is only guaranteed to be communicated between the host and devices at coarse grained synchronization points. Such synchronization is not suited for communication with persistently running kernels, for asynchronous updates to or from executing kernels, or for some low latency applications.