On the Efficiency Gains of Using Disaggregated Hardware to Build Warehouse-Scale Clusters
Giovanni Farias, Francisco Brasileiro, Raquel Borges de Freitas Lopes, Marcus Carvalho, Fábio Morais, Daniel Turull · 2017
Efficiently scheduling the workloads that are submitted to warehouse-scale clusters is not a trivial task. In these systems, the scheduler needs to deal with the heterogeneity in both the jobs that compose the workload, as well as in the servers that comprise the clusters. Moreover, placement constraints that either prevent or force jobs to be allocated in particular servers, makes the scheduler's task even harder. A number of strategies have been proposed to increase the efficiency of these schedulers, however, all of them assume the nowadays prevalent server-based architecture to build clusters. In this paper we assess the possible efficiency gains that can be attained considering that the underlying infrastructure is based on a disaggregated hardware (DH) architecture. This novel paradigm allows the dynamic assembling of logical servers from pools of system-widere sources, providing a way to shape the computing infrastructure while the workload is being allocated to it. Our simulation results, fed with publicly available data from relevant production systems, show that, on average, 5% more CPU demand, and 6% more RAM demand can be allocated, when we compare the fraction of a workload that a state-of-the-art scheduler is able to schedule on a server-based infrastructure with the one that it can allocate on an equivalent DH-based infrastructure. Moreover, in addition to allocating a larger workload, it can do so using less resources. For the workload we studied, on average, 8% of RAM capacity may be kept unused in the resource pools, where it can be switched off to save energy. Although at first sight these numbers might seem small, given the size of these systems, even small percentage improvements can lead to a very large economical impact.