STAR: a fault-tolerant system for distributed applications

Pierre Sens, Bertil Folliot · 2002

The paper presents a fault-tolerant manager for distributed applications. This manager provides an efficient recovery of hosts' failures on networks of workstations. An independent checkpointing is used to automatically recover application processes affected by host failures. Domino-effects are avoided by means of message logging and file versions management. STAR provides an efficient software failure detection by structuring hosts in a logical ring. Performance measurements in a real environment show the interest and the limits of our system.>

Read the paper · More papers on PaperTik