Hardware/software co-design and compiler techniques for efficient hardware acceleration of dense linear algebra kernels and machine learning applications:

Nícolas Bohm Agostini · 2024

Today's linear algebra and machine learning applications (ML) continue to grow in size and complexity, placing rapidly increasing demands on the underlying hardware and software systems. To address these issues, hardware designers have proposed using custom accelerators explicitly designed for these demanding workloads. However, realizing the full potential of custom accelerators requires effective hardware/software (HW/SW) co-design strategies that enable seamless integration and optimization across both domains, which is currently a major challenge in accelerator development. This thesis presents an integrated solution to facilitate HW/SW accelerator design. We also address issues in accelerator deployment, enabling rapid prototyping, integrated benchmarking, and comprehensive performance analysis of custom accelerators. In this thesis, we demonstrate the value of a lightweight system modeling library integrated into the build/execution environment of a production ML runtime engine, leveraging TensorFlow Lite for deployment. We explore custom accelerators performance by analyzing the trade-offs between various design choices, including parameter settings and accelerator configurations, from small to large scale, to identify the most effective combinations for specific workloads. Secondly, we employ the Multi-Level Intermediate Representation (MLIR) compiler framework to automatically partition host code from accelerator code, pre-optimizing the latter for improved high-level synthesis (HLS) designs and high-quality accelerated kernels. Lastly, we present compiler extensions to automate the generation and optimization of communication between the host CPU and AXI-based accelerators. We present novel solutions that enable more efficient and effective design space exploration, optimization, and deployment of custom accelerators. The utility of these approaches is demonstrated through experiments with specific accelerator designs and key linear algebra and ML workloads. Most importantly, these solutions empower high-level language users, such as domain scientists, to participate in the design of new accelerator features.--Author's abstract

Read the paper · More papers on PaperTik