Pat Helland on Failure and Resilience in Distributed Systems
Edaena Salinas · IEEE Software · 2019
Pat Helland of SalesForce talks with show-host Edaena Salinas about failure and resilience in distributed systems. The interview covers the growth of distributed computing on the public cloud, the origin of failure in the manufacturing process of hardware, the lifecycle of a server, why servers need to be replaced every three years, self-healing systems, replication, and the consistency, availability, and partition-tolerance (CAP) theorem. Sections that were not transcribed, due to space limits, cover how mutability makes life harder, append-only data, how electricity costs dominate, “cattle versus pets,” the rise of DevOps automation, testing, testing in production, and monitoring.