Distributed File System Multilevel Fault-Tolerant High Availability Mechanism

Yufeng Shi, Jinxin Zuo, Yihong Guo, Yueming Lu · 2020

To solve the single point of failure (SPoF), we designed a distributed file system multilevel fault-tolerant high availability mechanism and implemented the multilevel fault-tolerant file system (MFTFS). At the same time, we proposed a multilevel-raft election algorithm to automatically solve the split-brain and master state switching problems. By introducing a hot-standby master and a cold-standby master cluster, it is possible to quickly recover from failures to achieve high system availability when the primary master fails. The hot-standby master maintains a strict consistency of metadata in memory with the primary master, which facilitates rapid failover and replaces the master node to provide external services. The cold-standby master cluster backs up the metadata in the hard disk of the hot-standby master. At the same time, it acts as a third party to judge the status of the primary/hot-standby master, when the primary/hot-standby master is missing, it elects a new master, to avoid split-brain problems. Experimental results show that our mechanism can maintain the strict consistency of metadata in the primary/hot-standby memory, while making the mean failure recovery time less than 2s, and has an excellent performance in solving single-point failure and split-brain problems.

Read the paper · More papers on PaperTik