Fault Tolerance of the Application Manager in Vigne

Rajib Kumar Nath · 2008

Failure is a common phenomenon in distributed systems. As a system gets larger and complex, the numbers, sources and types of errors increase proportionately. To implement a system that works properly despite failures has been a major concern in distributed system research community for the past few decades. Researchers have been trying to enhance fault tolerance in every sector of computing such as web servers, storage, microprocessor, communication channel, application components etc. But ensuring fault tolerance in a large system like grid has been a very challenging and troublesome job. Grid is a collection of computing resources located in different administrative domain. Because of this large volume, it’s not easy to manage grid. The grid middleware was introduced to fulfill this need. Usually a user submits a job through grid middleware and gets the result back after the computation is over. Throughout the life time of the job, a grid middleware service component manages the job on behalf of the client. In the grid middleware Vigne [39, 32], developed by Paris project team in IRISA, Application Manager (AM) is the service component that supervises any application. It is the critical component of the system. That’s why we want it to keep running despite any failures or reconfiguration that can occur in the grid. The goal of our work is to find sollution to increase the availability of AM in Vigne. This report is organized as follows. Section 2 gives a short description of Vigne and presents the problem informally. Section 3 recognizes critical issues in a fault tolerant system. Then we describe the existing fault tolerance techniques and compared them in section 4 to select a suitable technique for replicating AM. Section 5 and section 6 make the necessary specifications to integrate the selected technique in Vigne. Section 7 addresses crucial issues regarding reconfiguration in highly dynamic environment. Related works are presented in section 8. Finally we present the conclusion and future work in section 9.

Read the paper · More papers on PaperTik