CUDA Memory Techniques for Matrix Multiplication on Quadro 4000

Tekesha Athil, Richard Christian, Yenumula B. Reddy · 2014

Today, the industry old adage of sequential processing is certainly no longer sufficient. The need for high performance computation is ever growing, even though certain problem sets remain within the realm of super high performance computing with applications such as weather forecasting, quantum physics and climate research to name a few. Within the commercial realm of computation, NVIDIA has proposed an architectural framework (NVIDIA CUDA) to harness the power of GPUs which before was only been utilized for Graphics Application like 3D games, but now recently, been used for certain types of high performance computation. In this paper, we will take a critical look at different performance techniques such as tiling, memory coalescing, perfecting, and loop unrolling, in trying to evaluate which method is the most efficient approach for our problem set (matrix operation of n size matrices).

Read the paper · More papers on PaperTik