Event management in computer networks
Yih-Farn Robin Chen, C. V. Ramamoorthy · 1987
In a network environment where events occur asynchronously on many computing nodes, the system state is changed rapidly in various ways. For clients to deal with these events, three network services are desirable: (1) Status service. Answering queries about the network state, (2) Notification service. Informing clients of interesting state changes, and (3)Action initiation service. Invoking computations on behalf of clients when the network state satisfies some predicates. A network event management system encapsulating these three services is proposed. We argue that a layered approach should be adopted to build such a system. The first layer is the communication subsystem layer. It provides primitives for inter-process communication and process control. The second layer is the clock synchronization layer. It guarantees that clocks of all non-faulty nodes are well synchronized. Network nodes communicate with each other to calculate the clock differences and adjust their clocks. The third layer is the status maintenance layer. A network server provides clients with a uniform access to a network status database, which implements a conceptual view of the network environment with a temporal dimension. The fourth layer is the event management layer. A network event manager accepts and registers service requests of the form $\{$event, action$\}$ from network clients. An event is specified as a predicate on the network state. When a registered event occurs, the corresponding action is fired. An action can be as simple as notifying the client, or as complex as initiating a remote process. The fifth layer is the application layer. With the services provided in the lower layers, clients can easily adapt to the changing network state. In particular, we show how the notification service simulates an active bulletin board and how the action initiation service supports the policy/mechanism separation paradigm. We then show how all the services contribute to the management of amoeba computations, a common form of large distributed computations. The lessons we learned in implementing each layer on a network of SUN-3/50 workstations are presented. Finally, we predict the impact of new technologies on the design of the network event management system.