A look at the OpenCL 2.0 execution model
Benedict R. Gaster · 2015
A popular approach to programming manycore GPUs is the Single Instruction Multiple Thread (SIMT) abstraction. SIMT has the benefit of presenting a "single thread" view, alleviating the complexity of explicitly vectorizing the source code. However, due to the SIMD nature of the underlying hardware it is often difficult to fully hide all aspects from the developer. An example of "leaks", is OpenCL's barrier, which requires all workitems (i.e. threads) to reach and execute the "same" barrier.