Efficient thread/page/parallelism autotuning for NUMA systems
Mihail Popov, Alexandra Jimborean, David Black-Schaffer · 2019
Current multi-socket systems have complex memory hierarchies with significant Non-Uniform Memory Access (NUMA) effects: memory performance depends on the location of the data and the thread. This complexity means that thread- and data-mappings have a significant impact on performance. However, it is hard to find efficient data mappings and thread configurations due to the complex interactions between applications and systems.