Operational performance metrics in a distributed system. Part II.

Robert L. Braddock, Michael R. Claunch, J. Walter Rainbolt · 1992

This paper presents a set of operational performance metrics to characterize operational distributed systems.Two high-level aggregale metrics, System Availability and System Throughput, are described, with subordinate metrics for each of the two groups identified.Definitions and goals are presented for each specific operational performance metric.Guidelines for interpreting these metrics are described and suggested corrective actions are provided.These metrics assist operations personnel in managing large, compiex distributed systems as integrated syslerns.1.0 performance of a distributed system and to provide for future capacity planning.Monitoring [he health and status of such a system requires the implementation of operational performance metrics.While operational performance metrics for centralized systems have been studied extensively [1, 2], little work has been reported on monitoring the performance of operational distributed systems in an integrated manner.This paper proposes a suite of operational performance metrics for distributed systems.Descriptions for each metric include the data being collected for the metric and the goals in monitoring the metric.Examples of the operational performance metrics are presented, further describing interpretation of the metric, corrective actions to be taken based on protrlcms identified by interpreting the metric, and definition of drc appropriate target audience.

Read the paper · More papers on PaperTik