Implementing halt on failure processors
R. Macdonald, Gholamali C. Shoja · 2002
The problem of detecting and masking failed processes in a distributed processing environment is considered. The authors propose a virtual halt on failure processor where replicated processes are used to achieve fault tolerance. Processor failures are detected and masked up to a certain limit. Once the threshold of permissible node failures is exceeded, the virtual processor reports the failure and halts. The authors contend that this is more practical and efficient than the generally assumed fail-stop processor. Results of an implementation in the REM (Remote Execution Manager) environment are presented.>