Limits on Thread-Level Speculative Parallelism in Embedded Applications

Mafijul Md. Islam, Alexander Busck, Mikael Engbom, Simji Lee, Michel Dubois, Per Stenström · 2007

As multi-core microprocessors are becoming widely adopted, the need to extract thread-level parallelism from sequential single-threaded applications in a seamless fashion increases. In this paper, we study the limits of performance speedup for embedded applications using parallelizing compilers on platforms with and without support for thread-level speculation. First and somewhat expected, only two out of ten applications from the consumer and telecom domains of the EEMBC suite could be automatically parallelized on multi-core architectures with no thread-level speculation (TLS) support. We systematically study the speedup obtained by parallelizing compiler technologies by factoring in the impact of the number of cores, thread decomposition strategies, and threadmanagement overhead. Overall, we have found that a TLS substrate is critical to uncover thread level parallelism and thread-management overhead must be low. On an eight-way multi-core system, it is possible to achieve a speedup of four, on average, for six out of the ten applications of EEMBC which we have analyzed. 1.

Read the paper · More papers on PaperTik