A fast deduplication scheme for stored data in distributed storage systems

Yuhang Long, Yingxun Fu · 2023

Data deduplication can effectively reduce data redundancy. However, its write performance is insufficient for existing storage systems, due to the additional calculation and I/O operations. In order to improve the deduplication speed in a distributed storage system, we propose FastDedup, a fast and effective deduplication scheme that focuses on the stored data. FastDedup improves deduplication speed through deduplication task distribution model and multi-container pool technology. Specifically, the deduplication task distribution model maintains the correctness for multiple deduplication nodes working simultaneously. The multi-container pool technology saves the operation time on the data merging stage. Evaluation results on three real backup datasets demonstrate that, compared to the unimproved technique, FastDedup increases deduplication throughput by 3.2% - 69.1%.

Read the paper · More papers on PaperTik