Mapping of a Machine Learning Algorithm Representation to Distributed Disaggregated FPGAs
Burkhard Ringlein · Zenodo (CERN European Organization for Nuclear Research) · 2023
Unofficial online version of the PhD thesis of Burkhard Ringlein submitted at the Faculty of Engineering of the Friedrich-Alexander-University Erlangen-Nürnberg in 2022 and published at Verlag Dr. Hut in 2023. The winding down of Moore’s law and the end of Dennard scaling have created a demand for specialized accelerators, including field-programmable gate arrays (FPGAs), in cloud and high-performance computing to fuel high demanding workloads, like machine learning models and artificial intelligence. At the same time, compute resources are increasingly consumed by cloud offerings to save costs and avoid the maintenance of sparsely utilized on-premise hardware. Despite their advantages in performance, adaptability, and energy-efficiency, FPGAs are not yet being deployed at scale in clouds and data centers due to their difficult tool support. Therefore, this thesis intends to contribute to a wider adoption of FPGAs in the future with improvements on three levels: First, a system architecture for managing a large number of disaggregated network-attached FPGAs in an efficient, flexible and scalable way is presented. To ensure the integrity of the infrastructure, partial reconfiguration is used to separate the non-privileged user logic from the privileged system logic. Furthermore, to create a scalable and agile cloud service, the management of all resources is built on the representational state transfer (REST) concept. Based on this, an FPGA design pattern is proposed to enable elastic microservices and function-as-a-service offerings. Using this system architecture, a provisioning time for new FPGA instances of below 7 seconds is demonstrated. The resulting combination of traditional CPU servers and FPGA nodes, which are all connected via the same network, leads to heterogeneous clusters for which no established programming model exists and which are hence cumbersome to use. Here, as a second level of this thesis, the proposed programming models for such clusters are revisited and it is argued that the Message Passing Interface (MPI) is suitable for programming CPU-FPGA clusters. Consequently, a one-click solution, called ZRLMPI, for compiling, optimizing, and deploying MPI applications on heterogeneous clusters is developed. The evaluation with clusters of up to 31 FPGAs shows a speedup of more than 25 times compared to clusters of CPUs. Finally, as a third level of this thesis, a framework for mapping deep neural network (DNN) models to distributed disaggregated FPGAs is developed. After assessing the current state-of-the-art of compilation frameworks for DNNs, the concepts of a meta-compiler and operation set architectures are presented and implemented. This meta-compiler, called DOSA, enables the evaluation, selection, and combination of existing but restricted DNN-to-FPGA tools to leverage previous research and to generate more efficient solutions. Moreover, DOSA can import DNNs represented in the community standard ONNX and implements model- and data-parallelism automatically, based on the performance targets and resource footprints provided by the user. Deploying a DNN on 9 FPGAs exhibits a speedup of more than 50 times compared to a CPU and 18 times compared to a GPU.