The Reliability Pillar

Ben Piper, David Clinton · 2019

This chapter focuses on the designing of Amazon Web Services (AWS) environment to tolerate resource failures so that the failure of an instance or even an entire availability zone doesn't cause your entire application to become unavailable. A common way of quantifying and expressing reliability is in terms of availability. Availability is the percentage of time an application is working as expected. Note that this measurement is subjective, and “working as expected” implies that you have a certain expectation of how the application should work. The level of availability your application can achieve depends in part on the availability of the AWS resources it uses, including networking, compute, and storage. Traditional applications are those written to run on and use the capabilities of traditional Linux or Windows servers. To deploy such an application on AWS, you'll need to run it on one or more EC2 instances.

Read the paper · More papers on PaperTik