PRISM: PRecision-Integrated Scalable Monitoring

Navendu Jain, Dmitry M. Kit, Prince Mahajan, Praveen Yalagandula, Mike Dahlin, Yin Zhang⋆ · 2008

This paper describes PRISM, a scalable monitoring service that makes imprecision a first-class abstraction. Exposing imprecision is essential for both correctness in the face of network and node failures and for scalability to large systems. PRISM quantifies imprecision along a three-dimensional vector: arithmetic imprecision (AI) and temporal imprecision (TI) balance precision against monitoring overhead while network imprecision (NI) addresses the challenge of providing consistency guarantees despite failures. Our implementation provides these metrics in a scalable way via (1) self-tuning of AI budgets to shift imprecision to where it is useful, (2) pipelining of TI delays to maximize batching of updates, and (3) dual-tree prefix aggregation which exploits regularities in our DHT’s topology to drastically reduce the cost of the active probing needed to maintain NI. PRISM’s careful management of imprecision qualitatively improves its capabilities. For example, by introducing a 10 % AI, PRISM’s PlanetLab monitoring service reduces network overheads by an order of magnitude compared to PlanetLab’s CoMon service, and by using NI metrics to automatically select the best aggregation results, PRISM reduces the observed worst-case inaccuracy of our measurements by nearly a factor of five.

Read the paper · More papers on PaperTik