Data placement in shared-virtual-memory multiprocessors with non- uniform memory access times

Jayashree Ramanathan · Michigan State University Libraries · 1992

Each processor in a multiprocessor usually has some local memory and also shares a global memory with the other processors. Such a memory organization is motivated by price/performance reasons, but results in a non-uniform memory access (NUMA) time. Many NUMA multiprocessors allow processes of an application to interact by means of a shared virtual memory because of the advantages of virtual memory and the ease of programming using the shared-memory model. In order to achieve an acceptable performance in these multiprocessors, an application's data should be judiciously placed in the memory hierarchy. The primary aim of this thesis is to identify the limitations of existing data placement techniques and to develop better methods of placing data. Existing data placement techniques include replication and migration, and are implemented in terms of blocks such as pages or cache lines. Using analysis and trace-driven simulations, we show that page-level replication needs to adapt to page reference patterns and hardware parameters. We also demonstrate that proper layout of data in the shared virtual memory simplifies page-level placement by reducing false sharing and simplifying reference patterns. These results also apply to other block-level placement strategies. Without help from the compiler or the applications programmer, adaptive strategies incur runtime overhead, and cannot handle false sharing and rapidly-changing reference patterns. These conclusions motivate our study of compiler-assisted data placement. Our approach uses compile-time objects containing data of the same variable type and similar reference patterns. We demonstrate how this approach can be incorporated in a compiler. We develop a compile-time object-creation scheme that assists block-level placement, and propose solutions to several related issues. We develop a scheme that assists object-level placement where the applications programmer specifies placement operations. For each scheme, we derive the compile-time overhead, and measure the performance of applications using experimental simulations. We demonstrate that there is significant performance improvement with compiler-assisted data placement, and that the performance is best when the compiler specifies all data placement operations. In summary, compiler-directed data placement is a promising approach to improve the performance of applications and the programmability of NUMA multiprocessors.

Read the paper · More papers on PaperTik