An evaluation of load sharing in locally distributed systems

Sayed A. Banawan · 1987

Load sharing has been the focus of a great deal of research as a means of enhancing the performance of distributed systems. This dissertation evaluates the potential benefits of load sharing when important system characteristics are varied. We start with a baseline model that consists of homogeneous, monolithic nodes executing independent jobs. Simulation results show that some obvious load measures, e.g., queue length and response ratio, can greatly improve system performance, while others, e.g., arrival rates and throughputs, have minor effect. In addition, we show that load balancing compares favorably with a greedy approach that optimizes the performance of individual jobs. The first variation to the baseline model allows heterogeneous nodes, i.e., nodes with different speeds. Using Markov decision theory we show that, except in highly heterogeneous systems, the queue length scaled by node speed yields an almost optimal policy. Furthermore, we demonstrated that performing load sharing while ignoring system heterogeneity can risk performance inferior to the no load sharing case. The second variation deals with cluster-based distributed systems. We show that most of the benefits can be realized by employing load sharing within each cluster alone. The additional benefit of inter-cluster load sharing is small and may not justify the potential overhead. The third variation considers multiple-resource nodes. The model captures systems with local I/O devices. We find that the performance gain from load sharing diminishes with node complexity regardless of routing, utilization, system size or the load measure used. However, should bottleneck resources exist or should nodes have different workload intensities a significant improvement becomes possible. Finally, we change the nature of workload to combine the two views of task allocation and job scheduling traditionally treated separately. Each job may have multiple concurrent modules. Experimental results with closed systems show that load sharing may degrade system performance since it exacerbates synchronization cost. A single overloaded node can delay many computations that have modules running on that node. In essence, the maximum speed achieved by a task is limited by the slowest node it uses. The same behavior was also observed in open systems where the number of jobs changes over time, based on simulation.

Read the paper · More papers on PaperTik