Performance Modeling of Shared Memory Multiple Issue Multicore Machines

Reshmi Mitra, Bharat S. Joshi, Arun Ravindran, Arindam Mukherjee, Ryan S. Adams · 2012

The process of developing optimal parallel applications is computationally expensive. The goal of this work is to design and validate a Markov chain based system-level performance prediction models to efficiently optimize parallel applications on shared memory multicore processors with coarse-grain thread level parallelism (TLP) like Intel Xeon Clover town. In Markov chain based throughput prediction model, the machine micro-architecture is represented by the different states and the allowable transitions. The program characteristics (such as cache misses, branch misprediction, division, denormalized computations and other large latency operations) are included using the failure probabilities of active and suspended threads. The improvement in performance is achieved by extracting information from running a representative data-set of the actual application. The model is validated with multiple benchmarks (electromagnetics application, parallel BZIP, FFT etc.) using VTune - Intel's performance analyzer. The average performance prediction error is less than 10%. The total run time for model is of the order of minutes (including VTune analyzer measurement timings), whereas the actual application is in terms of few hours.

Read the paper · More papers on PaperTik