DedupBench: A Benchmarking Tool for Data Chunking Techniques

Alan Liu, Abdelrahman Baba, Sreeharsha Udayashankar, Samer Al-Kiswany · 2023

Data deduplication is a technique for reducing storage space by identifying and eliminating redundant data. The division of files into chunks is one of the key steps in the deduplication process and directly impacts deduplication effectiveness. Despite the numerous algorithms available for chunking, there is a limited understanding of their strengths and weaknesses in virtual machine backup environments.We present DedupBench, a framework designed to assess the performance of different chunking algorithms for deduplication on user-specified data. DedupBench allows for the evaluation of chunking techniques by comparing their deduplication ratio and chunking throughput. DedupBench incorporates a generic design, allowing for the effortless integration of additional chunking techniques developed in the future.We evaluate four widely used chunking algorithms using a VM-based dataset with DedupBench. Our evaluation contrasts earlier studies and demonstrates that Asymmetric Extremum (AE) has the best deduplication efficiency for VM-based datasets among the tested algorithms, highlighting the need to evaluate chunking techniques on user-specified data before designing deduplication systems.

Read the paper · More papers on PaperTik