Efficiently enforcing strong memory ordering in GPUs

Abhayendra Singh, Shaizeen Aga, Satish Narayanasamy · 2015

GPU programming models such as CUDA and OpenCL are starting to adopt a weaker data-race-free (DRF-0) memory model, which does not guarantee any semantics for programs with data-races. Before standardizing the memory model interface for GPUs, it is imperative that we understand the tradeoffs of different memory models for these devices. While there is a rich memory model literature for CPUs, studies on architectural mechanisms and performance costs for enforcing memory ordering constraints in GPU accelerators have been lacking.

Read the paper · More papers on PaperTik