Dynamic Clustering-based Sharding in Distributed Deduplication Systems
Zhou Peng, Xiangyu Zou, Wen Xia · 2022
Witha massive upsurge in data, combining dedu-plication with distributed storage continuously suffer from a low deduplication ratio when providing the corresponding through-put. It is because distributed storage requires sharding data on different nodes, while global deduplication needs eliminating re-dundancies in a unified view. In this paper, we present clustering-based sharding method, D-Shard, in distributed deduplication storage systems that leads to a comparable deduplication effi-ciency on a single system while supporting a high throughput. To achieve that, we try to make that duplicates are located at the same node as much as possible with two key techniques. First, using Dynamic K-Means approach to cluster super-blocks, then extracting every cluster center feature as the anchor point for sharding; Second, Construct a secondary deduplication index based on the Compact Hamming Index. If a super-block mapping to the middle of two anchor points across nodes, then super-blocks adjacent to the feature are re-deduplicated in the cluster offline. Currently, preliminary results show that super-block clustering is convergent, and routing strategy based on anchor points can achieve a higher deduplication ratio compared to the state-of-the-art approach and the throughput of system has been greatly improved.