Improving GPU performance through instruction redistribution and diversification

Minglun Gong · 2018

As throughput-oriented accelerators, GPUs provide tremendous processing power by executing a massive number of threads in parallel. However, exploiting high degrees of thread-level parallelism (TLP) does not always translate to the peak performance that GPUs can offer, leaving the GPUs resources often under-utilized.

Read the paper · More papers on PaperTik