A Distributed Model of Computation for Reconfigurable Devices Based on a Streaming Architecture
Paolo Cretaro · 2019
FPGA-based devices are gaining attention in the High-Performance Computing research field because of their reconfigurability, which allows to implement highly optimized domain specific architectures. With the availability of High-Level Synthesis tools, it is far easier now to implement acceleration cores on FPGAs, but the standard CPU+accelerator paradigm, which is the dominant model for GPGPU, falls short exploiting one of the key feature these devices provide, which is the high-throughput and low-latency network capability. In this paper we show a model for leveraging these network capabilities inside a custom accelerator block, allowing to build a net of interconnected computing blocks, which are able to share data directly, without host participation.