Accelerating Dense Linear Algebra for GPUs, Multicores and Hybrid Architectures: an Autotuned and Algorithmic Approach

Rajib Kumar Nath · 2010

Dense linear algebra(DLA) is one of the most important softwares in high performance computing. It is also important for it’s wide usage in other application domains like machine learning, gaming, speech processing, image processing, etc. The introduction of new machines from vendor provides us opportunities to optimize DLA libraries for the new machines and thus exploit their power. Unfortunately the optimization phase is not straightforward all the time. The most important part of DLA libraries are it’s basic linear algebra subprograms(BLAS) kernels. The optimum code of a certain BLAS kernel in two different machines with different semiconductor process can be different even if they share the same features in terms of instruction set architecture, memory hierarchy and clock speed. It has become an tradition to optimize BLAS for upcoming machines. Vendors like Intel, AMD and IBM maintain highly optimized BLAS libraries targeting their own CPUs. In the GPU sector, NVIDIA is also providing CUBLAS for it’s accelerator cards like GTX280, Tesla2050. There has been few research in academia to optimize BLAS for GPUs. But the area is still new and presents numerous cases/opportunities for improvements. The existing BLAS for GPUs are not highly optimized for DLA algorithms. For example, vendors don’t have highly optimized BLAS for rectangular shaped problem size. Level 2 BLAS e.g. symmetric matrix matrix multiplication, which are very important for memory bound operations like tridiagonalization, performs poorly. In certain GPUs like GTX280 BLAS kernels have performance dips due to partition camping phenomenon in global memory modules. More importantly the existing BLASs are not optimized for generic

Read the paper · More papers on PaperTik