Activating Protection and Exercising Recovery Against Large-Scale Outages on the Cloud

Long Wang, HariGovind V. Ramasamy, Ruchi Mahindru, Richard E. Harper · 2016

Cloud computing provides rapid provisioning, convenient deployment, and simplified management of computing resources and applications with pay-as-you-go pricing models [1]. As more and more workloads are created on the cloud or migrated to the cloud for economic and flexibility reasons, it is important for developers, users, and service providers alike to understand the challenges, opportunities, complexities, and benefits of building dependable systems and applications on the cloud. In this tutorial, we give participants first-hand experience in protecting and recovering against large-scale cloud outages, e.g., failure of an entire cloud site. The tutorial is organized as a full-day activity and is designed to be hands-on. The content is targeted towards a broad audience of users, developers, practitioners, and researchers in the area of cloud computing. The content level is beginner-to-intermediate, suited for anyone with an undergraduate background in computer science (or equivalent) and basic programming skills. In the theory part of the tutorial, we introduce terminology, concepts, and metrics for providing resiliency on a cloud platform. We catalog factors that make building resilient applications on the cloud easy in some cases and particularly complicated in other cases. We present a reference architecture and a standard set of use cases for resiliency on the cloud. The bulk of the tutorial focuses on educating the audience with a series of hands-on exercises. We use an example set of cloud requirements as the starting point, and then guide the participants through the process of creating a protection and recovery plan. The plan covers details such as how to prioritize different workloads based on their criticality during recovery, what protection and recovery technologies should be used, and whether they should be used at the server level or application level. During the hands-on exercises, participants form teams or work individually to access a pre-created cloud virtual infrastructure and applications hosted on the IBM Softlayer cloud [2], which is geographically distributed across multiple continents. Replication and recovery orchestration form the backbone of many cloud resiliency solutions. We guide the participants through the entire life cycle of a cloud resiliency solution: 1) activation of protection on a set of workloads, 2) recovery of protected workloads upon a large-scale outage, 3) failback of protected workloads from the recovery site to the original site upon restoration of the original site, and 4) test of the implemented protection and recovery solution to ensure the implementation conforms to the requirements. Using a real-world orchestration technology, participants activate protection against outages at multiple levels of the cloud stack, orchestrate recovery procedure for a simulated site-level outage, and orchestrate failback to the primary cloud site (simulating the reconstruction of that site). We perform exercises for protecting and recovering both servers and applications using different types of replication technologies. The hands-on exercises are tailored to enable audience members to gain a strong grasp of the practical challenges involved in cloud resiliency, e.g., determining recovery priorities based on business criticality, recovery groups, and coordinated recovery across multiple virtual machines constituting a business application. Through the exercises, we reinforce core design principles and design elements for building resilient cloud applications. We expose the participants to a wide range of protection and recovery options for achieving resiliency against site-level cloud outages. Cloud users can use this experience to make more informed decisions on protecting their workloads against large-scale failures. After the hands-on exercises, we conclude with a survey of commercial and academic solutions, emerging areas, and future research challenges in the area of cloud resiliency.

Read the paper · More papers on PaperTik