Thread Migration to Improve Synchronization Performance
Srinivas Sridharan, Brett Keck, Richard C. Murphy, Surendar Chandra, Peter Michael Kogge · 2006
A number of prior research efforts have investigated thread scheduling mechanisms to enable better reuse of data in a processor’s cache. We propose to exploit the locality of the critical section data by enforcing an affinity between locks and the processor that has cached the execution state of the critical section protected by that lock. We investigate the idea of migrating threads to the “lock hot” processor, enabling the threads to reuse the critical section data from the processor’s cache and release the lock faster for other threads. We argue that this mechanism should improve the scalability of performance for highly multithreaded scientific applications. We test our hypothesis on a 4-way Itanium2 SMP running the 2.6.9 Linux kernel. We modified the Linux 2.6 kernel’s O(1) scheduler using information from the Futex (Fast User-space muTEX) mechanism in order to implement our policy. Using synthetic micro-benchmarks, we show 10-90 % performance improvement in cpu cycles, L2 miss ratios and bus requests for applications that operate on significant amounts of data inside the critical section protected by locks. We also evaluate our policy for the SPLASH2 application suite.