Efficient Erasure-Coded Data Recovery Based on Machine Learning With a Low Level of Storage Overhead
Xiaobo Zhao, Bing Yang Wei, Ning Luo, Qian Chen, Yi Wu, Shudong Zhang, Lijuan Marissa Zhou · 2024
Distributed storage systems typically use erasure codes for fault tolerance to reduce storage overhead. However, the data repair process in erasure-coded systems can generate heavy I/O overhead. Existing methods typically increase redundancy to improve repair speed, but this approach results in substantial storage overhead. To enhance repair speed while reducing storage costs, this paper proposes a Machine Learning-based Adaptive Recovery (MLAR) method. Given an application’s access patterns for a file, MLAR uses an adaptive encoding model to calculate the optimal code for each file. When applying fault tolerance with the optimal code, lower-redundancy codes can achieve faster repair speeds than higher-redundancy codes. MLAR employs machine learning to predict file access patterns. Experimental results replaying real-world I/O workloads show that, MLAR reduces storage overhead by 12.8% and recovery time by 23.7%, compared to state-of-the-art methods.