Using integrated compiler and architectural techniques to handle data dependences for thread-level speculative parallelization on multi-core processors

James M. Tuck, Gregory T. Byrd, Liang Han · 2011

Thread-Level Speculation (TLS) is a promising technique for improving performance of serial codes on multi-cores by automatically extracting threads and running them in parallel. However, the power efficiency as well as the performance gain of TLS systems are reduced due to frequent data dependence violations. Handling cross-thread dependences in the right way is key to improving the efficiency of TLS. This thesis studies efficient dependence handling techniques. The first study targets a class of dependences with reduction-like patterns. Reduction variables are an important class of cross-thread dependences that can be parallelized by exploiting the associativity and commutativity of their operation. In this work, we define a class of shared variables called partial reduction variables (PRV). These variables either cannot be proven to be reductions at compile time or they may be accessed outside their reduction operation statements and thus violate the requirements of a reduction variable. We describe an algorithm that allows the compiler to detect PRVs, and we also discuss the necessary requirements to parallelize detected PRVs. Based on these requirements, we propose an implementation in a TLS system to parallelize PRVs that works by a combination of techniques at compile time and in the hardware. The compiler transforms the variable under the assumption that the reduction-like behavior proven statically will hold true at runtime. However, if a thread reads or updates the shared variable as a result of an alias or unlikely control path, a lightweight hardware mechanism will detect the access and synchronize it to ensure correct execution. The second study broadly targets dependence prediction and synchronization. Prior work targeted frequently occurring dependences. But, in our well optimized TLS system, we find violations are caused by many infrequently occurring dependences too. In this work, we propose a new technique that is able to synchronize both frequently and infrequently occurring dependences in irregular tasking patterns without introducing excess synchronization. First, we enlist the help of the compiler to find and mark store-load pairs (SLPs) that generate data dependences. In this way, no training is needed to identify SLPs and irregular patterns are handled. Second, a hint operation informs hardware of a possible pending write so that a load only synchronizes when a store to the same address may occur in the future. The compiler schedules the hint as early as possible to warn successor threads. Finally, our mechanism wakes up a load as soon as possible using a release operation. Release is scheduled just after a store and on every path leading away from the hint in which the store does not occur. Together, they form our proposal called HiRe. We implement our compiler analysis and transformation in GCC, and analyze their potentials on a set of SPEC CPU 2000 benchmarks. We find that supporting PRVs provides up to 46% performance gain over a highly optimized TLS system and on average 10.7% performance improvement. And a TLS system supporting HiRe suffers only 22% of the violations that occur in our base TLS system, and it cuts the instruction waste rate of TLS in half. Furthermore, it outperforms prior approaches we studied by 3%.

Read the paper · More papers on PaperTik