Load Balancing by Remote Execution of Short Processes on Linux Clusters

M. Kacer, Pavel Tvrdı́k · 2003

PC clusters typically operate under continuously changing conditions and load balancing is critically important for efficient utilization of their resources, for maximizing their performance, and for minimizing the process response times. The average process response time is usually considered the most important value for measuring the actual performance of a multitasking system. Clearly, load balancing is critical for clusters running long-term processes (run times of hours). The main motivation of this paper is to study the suitability of the load balancing of short-term CPU-intensive processes, e.g., compilers, compression utilities, etc. Such processes represent a large part of a typical workload of a Linux workstation and load balancing can improve the process response time considerably in such cases. This topic appears to be neglected in the literature. The load balancing can be implemented by process migration (when a running process is stopped, checkpointed, checkpoint data are transferred to another node, and the process is restarted there) or by remote execution (when a process is transferred to another node when it is started and has no allocated memory). While the former solution is more universal (the process can be transferred in any moment), the latter one causes less overhead, which is very important especially when short processes are considered. A load-balancing algorithm is composed of three basic parts: a transferring mechanism, which implements remote execution or process migration, a load-balancing strategy, which makes decisions which process is to be moved where, and an information policy, which collects the information about the cluster state and distributes it among all the nodes. In Linux, a new process is created by fork system call, which makes an exact copy of the parent process, includ-

Read the paper · More papers on PaperTik