White Paper: A Grid Monitoring Service Architecture (DRAFT)

Brian Tierney, Ruth A. Aydt, Dan Gunter, Warren Smith, Valerie Taylor, Rich Wolski, Martin Swany · 2001

Large distributed systems such as Computational Grids require a large amount of monitoring data for a variety of tasks such as fault detection, performance analysis, performance tuning, performance prediction, and scheduling. Ensuring that all necessary monitoring is turned on and the data is being collected can be a very tedious and error-prone task. In this paper we propose an architecture for a Grid Monitoring Service to automate the execution of monitoring sensors and the collection event data. 1.0 Introduction The ability to monitor and manage distributed computing components is critical for enabling high-performance distributed computing. Monitoring data is needed to determine the source of performance problems and to tune the system for better performance. Fault detection and recovery mechanisms need monitoring data to determine if a server is down, and whether to restart the server or redirect service requests elsewhere. A performance prediction service might use monitori...

Read the paper · More papers on PaperTik