Reliable Networked and Distributed Systems
Ravishankar K. Iyer, Zbigniew T. Kalbarczyk, Nithin M. Nakka · 2024
In a distributed or networked system, many interconnected computers can cooperate to achieve a common goal. This chapter provides some introductory material on system/failure modeling; then, presents general problems (briefly introduced below) whose solutions, in the form of algorithms or protocols, constitute the following building blocks for solving application-specific problems: agreement, broadcast/group communication, software-based replication, and atomic commit. In a dependable distributed system, the individual processes that compose a distributed application are often required to reach mutual agreement on a common value. The distributed computing paradigm in data centers has continued to evolve to support increased performance, reliability, and efficiency. Resource disaggregation is the latest emerging paradigm. In disaggregated data centers, resources such as computing, storage, and memory are decoupled from their physical limits, allowing for dynamic allocation based on real-time demands and workloads.