SampDedup: Sampling Prediction for Efficient Inline Data Deduplication on Non-volatile Memory

Ziyue Xu, Yichen Li, Ranzhe Deng, Liping Yi, Yusen Li, Gang Wang, Xiaoguang Liu · ACM Transactions on Architecture and Code Optimization · 2025

Data deduplication is an effective technique for reducing redundant data storage space in various storage systems. Generally, deduplication consists of four steps: chunking, fingerprinting, fingerprint lookup, and data management. Recently, Non-volatile Memory (NVM) as an emerging storage device has received widespread attention. Directly applying the deduplication technique on NVM for storage cost savings faces many challenges: (a) deduplication on NVM devices suffers from computation bottleneck instead of the I/O bottleneck faced by deduplication on traditional storage devices (such as HDD and SSD); (b) new fingerprint indexes and metadata are required to be re-designed to adapt to NVM characteristics; (c) inline deduplication on NVM is more sensitive to the latency. To solve these challenges, we propose a novel Samp ling prediction-based inline data Dedup lication method ( SampDedup ) on NVM devices. It aims to ensure high deduplication ratios while reducing computation costs and latency by optimizing data chunking , fingerprinting , and fingerprint lookup . (a) For data chunking , a sampling prediction-based chunking method ( SampChunk ) is proposed to leverage chunk similarity to distinguish duplicate chunks and skip them for chunking. This method can be easily integrated into most sliding-window based and non-window based CDC chunking algorithms. (b) For fingerprinting , the commonly used SHA-1 algorithm is further optimized to reduce the extra computational overhead introduced by SampChunk, and an asynchronous fingerprinting method is proposed to reduce the fingerprinting latency of unique chunks. (c) For fingerprint lookup , we design a header fingerprint index and metadata table for each data chunk constructed by SampChunk on NVM, and we use a fast-read buffer to replace the traditional slow LRU cache to improve search efficiency. Experiments on four real-world datasets demonstrate that SampDedup consistently presents high inline data deduplication ratios on NVM with different workloads and data partitioning algorithms while saving more than 90% chunking time compared with state-of-the-art deduplication baselines.

Read the paper · More papers on PaperTik