A performance model for unified parallel c
Steven Seidel, Zhang Zhang · 2006
This research is a performance-centric investigation of Unified Parallel C (UPC), a parallel programming language that belongs to the Partitioned Global Address Space (PGAS) language family. The objective is to develop a performance modeling methodology that targets UPC but can be generalized for other PGAS languages. The performance modeling methodology is based on platform characterization and program characterization, achieved through shared memory benchmarking and static code analysis, respectively. Models built using this methodology can predict the performance of simple UPC application kernels with relative errors below 15%. Besides performance prediction, this work provides a framework based on shared memory benchmarking and code analysis for platform evaluation and compiler/runtime optimization studies. A few platforms are evaluated in terms of their fitness to UPC computing. Some optimization techniques, such as remote reference caching, is studied using this framework. A UPC implementation, MuPC, is developed along with the performance study. MuPC consists of a UPC-to-C translator built upon a modified version of the EDG C/C++ front end and a runtime system built upon MPI and POSIX threads. MuPC performance features include a runtime software cache for remote accesses and low latency access to shared memory with affinity to the issuing thread. In this research, MuPC serves as a platform that facilitates the development, testing, and validation of performance microbenchmarks and optimization techniques.