High throughput GPU polar decoder
Yijin Li, Rongke Liu · 2016
In this paper, a polar decoder with simplified decoding algorithm, optimized thread execution and memory access on a Graphics Processing Unit (GPU) is proposed. Four main aspects are optimized to improve the throughput of the polar decoder. Firstly, decoding information is arranged to ensure global memory coalesced access. Secondly, to avoid extra computation and access delay of frozen set, semi-unrolled architecture is adopted. Finally, the branch divergence is optimized to decrease the branch delay. The experiments are carried out on Nvidia Tesla K20 platform. Results demonstrate that the proposed decoder achieves about 700Mbps.