Flexible architectural support for fine-grain scheduling
Daniel Sánchez, Richard M. Yoo, Christos Kozyrakis · 2010
To make efficient use of CMPs with tens to hundreds of cores, it is often necessary toexploit fine-grain parallelism. However, managing tasks of a few thousand instructions is particularly challenging, as the runtime must ensure load balance without compromisinglocalityandintroducingsmalloverheads.Software-onlyschedulers can implement various scheduling algorithms that match the characteristics of different applications and programming models, but suffersignificant overheads astheysynchronize andcommunicatetaskinformationoverthedeepcachehierarchyofalarge-scale CMP. To reduce these costs, hardware-only schedulers like Carbon, which implement task queuing and scheduling in hardware, have been proposed. However, a hardware-only solution fixes the scheduling algorithm and leavesno room for other uses of the custom hardware. This paper presents a combined hardware-software approach to build fine-grain schedulers that retain the flexibility of software schedulers while being as fast andscalable as hardware ones. Wepropose asynchronous directmessages (ADM),asimplearchitectural extension that provides direct exchange of asynchronous, shortmessages betweenthreads intheCMPwithoutgoingthrough the memory hierarchy. ADMissufficient toimplement a familyof novel, software-mostly schedulers that rely on low-overhead messaging to efficiently coordinate scheduling and transfer task information. These schedulers matchand often exceed the performance and scalability of Carbon when using the same scheduling algorithm. When the ADM runtime tailors its scheduling algorithm to application characteristics,it outperforms Carbonby upto70%.