Data Deduplication Cluster Based on Similarity-Locality Approach
Xingyu Zhang, Jian Zhang · 2013
Human beings have entered the big data era, and the growing data bring huge challenges for data storage. Existing deduplication methods do not work adequately in many situations. Recently, the data deduplication cluster has become an important need of most commercial and research backup systems. Data deduplication cluster becomes popular in storage system for data backup and archiving. Many researchers focus on deduplication cluster by which to reduce more redundant data. Especially block level deduplication cluster becomes popular. It is concerned to have two challenges: the chunk-lookup disk bottleneck problem and the data routing problem. A new solution is proposed for chunk-lookup disk bottleneck in our paper. The Approach of combining similarity with locality is applied to the deduplication cluster. At the same time, the bloom filter algorithm storing fingerprint is used to find more duplicate data between nodes. The system architecture and the details are provided. Finally, the experiment shows the system has a good performance.