Locality Aware DAG-Scheduling for LU-Decomposition
Tobias Maier, Peter W. Sanders, Jochen Speck · 2015
Modern computers have deepening memory hierarchies with multiple levels of (partially shared) caches and non-uniform memory access (NUMA). This makes it increasingly difficult and important to schedule computations in such a way that expensive memory accesses are avoided.In this paper we are choosing LU-decomposition for a case study since its use in the famous LINPACK benchmark means that highly tuned codes are already available. Our approach is to perform the very same computations as a leading implementation (PLASMA) but to schedule them in a more locality aware way. In particular, we better take into account when independent subtasks share the same input data and we explicitly address NUMA-effects coordinating memory layout and task scheduling. These measures lead to up to 36 % performance improvement compared to PLASMA.