Coscheduling in the multicore era
Jan Hendrik Schönherr · DepositOnce · 2019
In its most general definition, coscheduling refers to a deliberate simultaneous execution of certain tasks on multiple CPUs. The concepts of coscheduling and gang scheduling have been introduced in the early eighties and nineties, respectively. At that time, the first massively parallel systems were developed and became more widely used. The guarantee of simultaneous execution of certain tasks allows to make use of fine-grained synchronization efficiently: it is impossible for a currently executing task to wait for another task that does not make progress. Compared to batch processing and partitioning in space, coscheduling allows to preempt a running parallel application in favor of another more important application. This context switch at application level offers more flexibility for schedulers. Since then, there have been thorough changes in the computing landscape: (non-parallel) personal computers have penetrated the market, and with it, parallelism had become a second-class citizen. Parallelism has only recently become important again, after single-thread performance could not be increased by the usual margins anymore. However, parallelism today differs from parallelism back then, mostly because of the following two reasons: 1. The reintroduction of parallel systems happened not overnight but as a gradual process. Due to this, existing software was adapted to run on parallel systems instead of being rewritten. This adapted software handles contemporary multicore systems as if they were distributed systems instead of the parallel systems they are: current operating systems have no concept of a parallel application; they do not know about coscheduling or other management approaches, hampering real parallel applications. To make matter worse, budding software developers (outside of the HPC niche) are seldom taught the subtleties of parallel software development. 2. The properties of today’s multicore architectures differ substantially from early parallel systems and clusters. For this reason it is not always possible to reuse once valid solutions. In particular, CPUs in today’s system have to share a multitude of resource, such as memory bandwidth of caches. This results in resource contention with a more or less noticeable impact on performance of individual tasks. In most cases research suggests to avoid or reduce resource contentions by grouping tasks skillfully – which is basically a form of coscheduling. However, an integration of these research ideas into existing systems is impractical more often than not, because individual ideas – while solving their specific use case – are difficult to combine with their rather narrow focus. This thesis is ports the concept coscheduling to contemporary multicore architectures. Advantages and disadvantages known from other parallel architectures are reevaluated, and a catalog of use cases – old and new – is compiled. The combined requirements of these use cases form the basis for a versatile coscheduling model. A flexible coscheduling approach is devised, which measures up to today’s expectations of application and operating system developers alike. In particular, a method is included that allows to integrate coscheduling functionality into existing general purpose schedulers without changing their characteristic traits substantially. Newly enabled management opportunities within the operating system are discussed and their influence on the design of upcoming parallel applications and operating systems are explored.