Handling cascading failures: the case for topology-aware fault-tolerance

Soila Pertet, Priya Lakshmi Narasimhan · Hot Topics in System Dependability · 2005

Large distributed systems contain multiple components that can interact in sometimes unforeseen and complicated ways; this emergent vulnerability of complexity increases the likelihood of cascading failures that might result in widespread disruption. Our research explores whether we can exploit the knowledge of the system's topology, the application's interconnections and the application's normal fault-free behavior to build proactive fault-tolerance techniques that could curb the spread of cascading failures and enable faster system-wide recovery. We seek to characterize what the topology knowledge would entail, quantify the benefits of our approach and understand the associated trade offs.

Read the paper · More papers on PaperTik