Failover Pattern with a Self-Healing Mechanism for High Availability Cloud Solutions

Alexander Stanik, Mareike Höger, Odej Kao · 2013

Cloud computing has already been adopted in a broad range of application domains and has become an established building block in IT landscapes. During the process of cloud middleware development, the companies have focused mainly on the high availability of data and end-user services, but unfortunately neglected the availability of middleware components. Therefore failures of the middleware components itself usually leads to a partial or even total blackout of the cloud. In this paper, we present the design and implementation of a novel scalable and highly available multi-master pattern for cloud middlewares. In contrast to existing Infrastructure-as-a-Service cloud management frameworks, which are usually designed in a centralized tree topology composed in a three-tiered master worker architecture, we introduce a concept for a multi tree with all tree roots connected in a fully connected mesh topology. In this architecture user requests are load balanced over multiple failover servers. Furthermore, our concept includes an automatic self-healing mechanism for worker nodes of each tree.

Read the paper · More papers on PaperTik