Replay: A Model-Based Service for Supporting Transparent Cluster Analysis Tools
Diwakar Krishnamurthy, Cameron Kiddle, Jerry Rolia, Rob Simmonds · 2006
Grid computing environments typically federate heterogeneous resource clusters belonging to several organizations. To fully realize the promise of a grid environment, it is necessary to support tools that help obtain insights into the behaviour of individual clusters. This paper describes a cluster service called Replay that simplifies the development and maintenance of such tools. The service provides a common model-based interface for obtaining current and historical information about a cluster. Replay can manage multiple views which allows tools to obtain information about an existing cluster as well as information that shows how a cluster might have behaved under alternate configurations and workloads. The model-based interface allows tools to be ported to different clusters with little effort. Furthermore, Replay uses different mechanisms to manage information that typically changes infrequently and information that can change in a more dynamic, continuous manner allowing it to handle information more efficiently than existing services that provide similar functionality. The paper presents a job analysis tool to illustrate the utility of the service