Automatically exploiting the memory hierarchy of GPUs through just-in-time compilation
Michail Papadimitriou, Juan José Fumero, Athanasios Stratikopoulos, Christos Kotselidis · 2021
Although Graphics Processing Units (GPUs) have become pervasive for data-parallel workloads, the efficient exploitation of their tiered memory hierarchy requires explicit programming. The efficient utilization of different GPU memory tiers can yield higher performance at the expense of programmability since developers must have extended knowledge of the architectural details in order to utilize them.