Interleaved Execution of Approximated CUDA Kernels in Iterative Applications

Gabriel Freytag, Cristiano A. Künas, Paolo Rech, Philippe O. A. Navaux · 2024

Fine-tuning the floating-point precision of arithmetic operations in applications can be extremely challenging and time-consuming, especially in iterative applications where the output of one iteration serves as the input for the subsequent iteration. Consequently, the accuracy loss can be magnified throughout the execution. Therefore, we propose an alternative approach based on the interleaved execution of multiple approximated kernel versions designed with different precision levels. We demonstrate that creating an interleaved execution configuration of multiple CUDA kernel versions based on their accuracy loss profiles enables us to enhance performance, improve energy efficiency, and manage the accuracy loss of scientific simulation applications in various Target Output Quality (TOQ) scenarios. For a TOQ loss of approximately 3 %, we achieve a speedup of up to 1. 7x and reduce energy consumption by nearly 40 %.

Read the paper · More papers on PaperTik