Tuning compiler optimizations for simultaneous multithreading

Jack L. Lo, Susan J. Eggers, Henry M. Levy, Sujay S. Parekh, Dean Michael Tullsen · 1997

Compiler optimizations are often driven by specific assumptions about the underlying architecture and imple-mentation of the target machine. For example, when tar-geting shared-memory multiprocessors, parallel programs are compiled to minimize sharing, in order to decrease high-cost, inter-processor communication. This paper reexamines several compiler optimizations in the context of simultaneous multithreading (SMT), a processor architecture that issues instructions from mufti-ple threads to the functional units each cycle. Unhke shared-memory multiprocessors, SMT provides and bene-fits from fine-grained sharing of processor and memory system resources; unlike current uniprocessors, SMT exposes and bene$ts from inter-thread instruction-level parallelism when hiding latencies. Therefore, optimiza-tions that are appropriate for these conventional machines may be inappropriate for SMi’I We revisit three optimiza-tions in this light: loop-iteration scheduling, software speculative execution, and loop tiling. Our results show that all three optimizations should be applied differently in the context of SMT architectures: threads should be paral-lelized with a cyclic, rather than a blocked algorithm; non-loop programs should not be software speculated, and compilers no longer need to be concerned about precisely sizing tiles to match cache sizes. By following these new guidelines, compilers can generate code that improves the pelformance of programs executing on SMT machines. 1

Read the paper · More papers on PaperTik