Tolerating Correlated Failures in Wide-Area Monitoring Services
Suman Nath, Haifeng Yu, Phillip B. Gibbons, Srinivasan Seshan · 2004
Recently, there has been increasing deployment of distributed infrastructure and applications such as Akamai, PlanetLab and DHTs. A highly-available monitoring service for these systems is vital to ensuring their smooth operation. In this setting, a key challenge to high availability is the presence of correlated failures, not addressed by previous monitoring services. This paper presents our approach to achieving high availability in IRISLOG, a customizable wide-area monitoring service. IRISLOG incorporates a replication design based on Signed Quorum Systems and a load-balancing design based on a novel XML database-fragmentation algorithm to achieve our availability goals. Using a combination of simulation under a tunable correlated failure model and results derived from a real world deployment on PlanetLab, we show that IRISLOG is able to achieve an availability level significantly higher than previous techniques.