High-speed volume ray casting with CUDA
Lukas Marsalek, Armin Hauber, Philipp Slusallek · 2008
By creating an optimized CUDA implementation of volume ray casting with pre-integration and early ray termination, we present a proof-of-concept that the enhanced flexibility of programming GPUs in C dialect does not come at a performance hit. Rather it enables low-level access to the hardware, outperforming optimized shader implementations by 12% to 36%, and naïve CUDA implementations by 68% to 114%. It also scales well across different hardware (8600 GTS and 8800 GT), as reducing the processing power by a factor of 3.5, reduces the rendering speed only by a factor of 1.3 to 1.9.