Survivability management architecture for very large distributed systems

John C. Knight, Jonathan C. Rowanhill · 2004

Our society is critically reliant upon national distributed systems such as electrical power grids, freight rail systems, and communication networks. Unfortunately these systems are vulnerable to damage and attacks. Traditional engineering practices can only partially mitigate these vulnerabilities, due to the scale, dynamics, and exposure of infrastructure systems. Engineers now propose designing systems to survive in more hostile environments. Survivable systems would operate as specified under a wider range of potential conditions, including wide-scale natural disasters and malicious attacks. Willow is a reactive survivability architecture supporting rapid reaction to non-local faults. Non-local faults are those involving distributed system states around which no clear, simple boundary may be drawn. Key to Willow's design is the ability to rapidly respond to detected non-local faults with non-local adaptations. We assert that a new model of management architecture—one based on high-level, loosely coupled management relationships—enables reactive survivability over non-local faults. The management model we introduce allows the rapid, heterogeneous distribution of control over very large systems. Control distribution is flexible and dynamic, so that it reflects conditions “on the ground.” Therein large-scale system adaptation and posturing is achieved. We synthesize and develop a loosely coupled management model. Our implementation, called ANDREA, includes novel technology contributions. The most important is Selective Notification with Reply, a scalable, symmetrically addressed, and loosely coupled communications service. We present both analytical and experimental analyses, with the latter based on a full implementation of survivability management service.

Read the paper · More papers on PaperTik