A Fast and Generic GPU-Based Parallel Reduction Implementation
Walid Abdala Rfaei Jradi, Hugo Alexandre Dantas do Nascimento, Wellington Santos Martins · 2018
Reduction operations are extensively employed in many computational problems when a finite set of numeric elements are combined into a single value using for this a combining function. A parallel reduction, in turn, is the operation concurrently performed when multiple execution units are avai\-lable. The present work depicts a GPU-based parallel approach for it, which employs techniques like loop unrolling, persistent threads and algebraic expressions to avoid thread divergence, able to surpass the methods currently in use. Experiments conducted to evaluate the approach show that the strategy performs efficiently on both AMD and NVidia's hardware platforms, as well as using OpenCL and CUDA, making it portable.