Automatic parallelization of fine-grained meta-functions on a chip multiprocessor

Sanghoon Lee, James M. Tuck · 2011

Due to the importance of reliability and security, prior studies have proposed inlining meta-functions into applica-tions for detecting bugs and security vulnerabilities. How-ever, because these software techniques add frequent, fine-grained instrumentation to programs, they often incur large runtime overheads. In this work, we consider an automatic thread extraction technique for removing these fine-grained checks from a main application and scheduling them on helper threads. In this way, we can leverage the resources available on a CMP to reduce the latency and overhead of fine-grained checking codes. Our parallelization strategy automatically extracts meta-functions from the main application and executes them in customized helper threads — threads constructed to mirror relevant fragments of the main program’s behavior in or-der to keep communication and overhead low. To get good performance, we consider optimizations that reduce com-munication and balance work among many threads. We evaluate our parallelization strategy on Mudflap, a pointer-use checking tool in GCC. To show the benefits of our technique, we compare it to a manually parallelized version of Mudflap. We run our experiments on an ar-chitectural simulator with support for fast queueing oper-ations. On a subset of SPECint 2000, our automatically parallelized code is only 29 % slower, on average, than the manually parallelized version on a simulated 8-core system. Furthermore, two applications achieve better speedups us-ing our algorithms than with the manual approach. Also, our approach introduces very little overhead in the main program — it is kept under 100%, which is more than a 5.3 × reduction compared to serial Mudflap. 1.

Read the paper · More papers on PaperTik