PFCG: Improving the Restore Performance of Package Datasets in Deduplication Systems
Chunxue Zuo, Wang Sheng Fang, Ping Sheng Huang, Yuchong Hu, Dan Feng, Yucheng Zhang · 2018
Data deduplication, a lossless data compression technique, has been widely deployed in backup systems to save storage space. However, data fragmentation introduced by data deduplication seriously degrades restore performance. To alleviate the fragmentation problem, rewriting algorithms such as CBR, CAP, and HAR have been proposed. In certain backup datasets containing packed files like UNIX tar files, we have observed that a large amount of rewritten fragmented chunks tend to remain fragmented in following backups and thus get rewritten repeatedly in consecutive backups. Such repeatedly rewritten chunks are referred to as persistent fragmented chunks (PFCs), and they severely impact the efficiency of rewriting algorithms and restore performance. In this paper, we propose PFCG, an efficient scheme to enhance the efficiency of rewriting algorithms and restore performance. The central idea of PFCG is to identify and differentiate PFCs from regular fragmented chunks, i.e., non-PFCs, and store these two types of fragmented chunks into separate containers during backups to prevent the proliferation of PFCs in following backups. Our experimental results based on six real-world datasets demonstrate that PFCG significantly improves the restore performance by 21% to 47% over the state-of-the-art rewriting algorithms, without sacrificing deduplication efficiency.