Monitoring of distributed systems

Young-Chul Shim, C. V. Ramamoorthy · 2002

Distributed systems offer opportunities for attaining high performance, fault-tolerance, information sharing, resource sharing, etc. But we cannot benefit from these potential advantages without suitable management functions such as performance management, fault management, security management, etc. Underlying all these management functions is the monitoring of the distributed system. Monitoring consists of collecting information from the system and detecting particular events and states using the collected information. These events and states can be symptoms for performance degradations, erroneous functions, suspicious activities, etc. and are subject to further analysis. Detecting events and states requires a specification language and an efficient detection algorithm. The authors introduce an event/state specification language based on classical temporal logic and a detection algorithm which is a modification of the RETE algorithm for OPS5 rule-based language. They also compare their language with other specification languages.>

Read the paper · More papers on PaperTik