Generating CUDA code at runtime: specializing accelerator code to runtime data
Tristan Perryman, Paul H. J. Kelly, Anton Lokhmotov, Tony Field · 2008
This abstract presents preliminary results from exploring the idea of generating code for a GPU accelerator at run-time. We show that this can lead to performance improvements due to specialization: we specialize the GPU code to the particular run-time data. We illustrate the prototype tool with a ray tracing example, where we achieve 10%–30 % speedup due to specialisation to the scene being rendered. The prototype tool is based on our Taskgraph library for runtime code generation in C++; this supports runtime generation of both the CUDA kernel itself, and the host code for transferring parameters and results. The tool gives the programmer full, dynamic control over CUDA resource management and storage assignment. 1.