High‐Performance Optimizations on Tiled Manycore Embedded Systems: A Matrix Multiplication Case Study*

Arslan Munir, Ross Gordon, Sanjay Ranka · 2016

This chapter gives an overview of architectural features of contemporary tiled manycore architecture (TMA), and summarizes work related to performance analysis on multicore architectures, parallelized matrix multiplication (MM) algorithms, cache blocking, and TMAs. It defines parallel computing metrics for TMAs and outlines the dense MM algorithms considered in the case study. Parallel computing metrics quantify the performance and performance per watt of parallel architectures, such as TMAs, and enable architectural comparisons. The chapter provides code snippets of the MM algorithms for Tilera's TILEPro64. Performance optimizations for TMAs including platform optimizations and compiler-based optimizations are discussed. The compiler-based optimizations include scalar optimizations, function inlining, alias analysis, loop unrolling, loop-nest optimizations, software pipelining, and feedback-based optimizations. The chapter presents the performance optimization results for the MM case study on Tilera's TMAs, with a focus on the TILEPro64. Finally, it summarizes insights obtained from this study.

Read the paper · More papers on PaperTik