Remote backup system based on data de-duplication

Liu Xiaojie · Jisuanji gongcheng yu sheji · 2012

To the problem that a large number of redundant data caused inefficient backup and storage waste in traditional remote backup,a remote backup system based on data de-duplication is designed and implemented.Backup files are divided into variable-length chunks based on Rabin fingerprint of contents.Chunks' information is sent to backup centre where duplicate chunks are sought by using Google Bigtable and Leveldb index algorithm along with bloom filter.Finally,it only transmitted and stored unique chunks.Experimental results show that,it can remove duplicate data effectively to backup similar data sets.Compared with Rsync backup,it has less network flow when it does a incremental backup which has small incremental data.

Read the paper · More papers on PaperTik