Hardware-level thread migration in a 110-core shared-memory multiprocessor
Mieszko Lis, Keun Sup Shim, Brandon Cho, Ilia Lebedev, Srinivas Devadas · 2013
Advantages - significantly reduces traffic on high-locality workloads up to 14x reduction in traffic in some benchmarks - simple to implement and verify (indep. of core count, no transient states) - decentralized & trivially scalable (only # core ID bits, addr ↔ core mapping) Challenges - workloads should be optimized with memory model in mind (like allocating data on cache line boundaries but more coarse-grained) - automatically mapping allocation over cores not a trivial problem Opportunities - fine-grained migration is an enabling technology - since it's cheap and responsive, can be used for almost anything - e.g., if only some cores have FPUs, migrate to access FPU.