On efficiently tolerating general failures in autonomous decentralized multiserver systems

I‐Ling Yen, Farokh Bastani · 2002

We consider a multiserver system consisting of a set of servers that provide some service to a set of clients by accessing some shared objects. The goal is to provide reliable service in spite of client or server failures such that the overhead during normal operating periods is low. We consider a relatively general fault model where a faulty processor can write spurious data for a period of time before it is detected and removed from the system. We first develop a solution in an autonomous physical world of clients and servers. The basic approach is to divide the servers into groups such that each server has some limitations which prevent it from arbitrarily damaging the system. This solution is then mapped to a distributed system of processor and memory units followed by an assessment of its performance.>

Read the paper · More papers on PaperTik