Thinking about Availability in Large Service Infrastructures
Jeffrey C. Mogul, Rebecca Isaacs, Brent B. Welch · 2017
Our company has learned to design and operate planetaryscale services with reasonably high availability.Historically, these have been Software as a Service (SaaS) systems (search, YouTube, GMail, etc.), implemented as scale-out distributed systems that could tolerate all sorts of failures in lower layers, through the use of traditional techniques such as replication, distributed consensus algorithms (e.g.Paxos), and transactions,