High Availability through Distributed Control

Christian Engelmann, Stephen L. Scott, George A. Geist · 2004

Cost-effective, flexible and efficient scientific simulations in cutting-edge research areas utilize huge high-end com-puting resources with thousands of processors. In the next five to ten years the number of processors in such computer systems will rise to tens of thousands, while scientific ap-plication running times are expected to increase further be-yond the Mean-Time-To-Interrupt (MTTI) of hardware and system software components. This paper describes the on-going research in heterogeneous adaptable reconfigurable networked systems (Harness) and its recent achievements in the area of high availability distributed virtual machine environments for parallel and distributed scientific comput-ing. It shows how a distributed control algorithm is able to steer a distributed virtual machine process in virtual syn-chrony while maintaining consistent replication for high availability. It briefly illustrates ongoing work in heteroge-neous reconfigurable communication frameworks and secu-rity mechanisms. The paper continues with a short overview of similar research in reliable group communication frame-works, fault-tolerant process groups and highly available distributed virtual processes. It closes with a brief discus-sion of possible future research directions. 1

Read the paper · More papers on PaperTik