The CEDPS troubleshooting architecture and deployment on the open science grid

Brian L. Tierney, Dan Gunter, Jennifer M. Schopf · Journal of Physics Conference Series · 2007

Tracking failures and poor performance across a widely distributed system of resources has proven challenging for many ongoing DOE applications. An example is the Open Science Grid (OSG) project, which currently experiences a roughly 15% job failure rate. This can be an issue not only for Grid computing but for anyone performing large-scale data transfers to remote machines because of the large number of interconnected components and services. As part of the Center for Enabling Distributed Petascale Science (CEDPS) project we have been building an infrastructure to work with current middleware and existing system tools to more easily track failures and discover anomalous behavior. This consists of a common logging format, the extension of syslog-ng for centralized collection of data, a data summarizer to more easily manage the volume of logging, and an anomaly detection system that can connect to a warning system when unexpected behaviors occur. We are currently working with OSG to deploy a prototype of the full system. The initial logs gathered will be used to extend the analysis tools and to increase the reliability of the services for the SciDAC end user community.

Read the paper · More papers on PaperTik