Weak Network Oriented Mobile Distributed Storage: A Hybrid Fault-Tolerance Scheme Based on Potential Replicas
Xinlei Wei, Yujuan Tan, Duo Liu, Daitao Wu, Yu Wu, Xianzhang Chen, Jian Li · 2022
Node failure is one of the most typical issues in distributed storage systems. The classic fault-tolerance methods can meet the fault tolerance needs of systems deployed in edge storage, 5G IoT, and other high-performance data centers. However, mobile distributed systems contend with inconsistent network signals and relatively low bandwidth in weak network environments. The traditional fault-tolerance schemes have difficulty ensuring distributed storage and data repair simultaneously, posing a considerable challenge to data storage reliability. To address this problem, we proposed a hybrid fault-tolerance scheme based on potential replicas, called HFPR, which trades off storage overhead and performance for weak network mobile distributed storage systems. HFPR first stores data with the erasure code redundancy and then gradually increases data redundancy by reserving potential replicas, enabling the system to disperse network transmission pressure and reduce network transmission. To maximize replica utilization and control the storage overhead of potential replicas, we design a node state prediction mechanism and a file lifecycle management mechanism. Evaluation results show that, compared with the traditional hybrid fault-tolerant scheme with the same storage overhead, HFPR achieves better replica-level repair efficiency in data repair and a reduced probability of degraded reads by approximately 50%. During data distribution, HFPR can minimize data write amplification by up to 1.1× and additional network transfers by up to 77.5%.