CACH-Dedup: Content Aware Clustered and Hierarchical Deduplication

Girum Dagnaw, Wang Hua, Ke Zhou · 2018

Distributed deduplication overcomes, to some extent, index-lookup disk bottleneck problem by dividing deduplication tasks among many nodes. However, the task of selecting these nodes is an important challenge because it could result in high communication cost and the storage node island effect problem. Moreover, intelligent data routing is required to exploit the peculiar nature of data from different applications which share insignificant amount of content. In this paper, we explore CACH-Dedup, a content aware clustered and hierarchical deduplication system, which exploits the negligibly small amount of content shared among chunks from different file types to create groups of files and storage nodes with out loss of deduplication effectiveness. It uses hierarchical deduplication to reduce the size of fingerprint indexes at the global level, where only files and big sized segments are deduplicated. It also makes advantage of locality first using the big sized segments deduplicated at the global level and second by routing a set of consecutive files together to one storage node. Furthermore, it exploits similarity by making use of similarity bloom filters of streams for stateful routing which results in duplicate elimination rate in a par with single node deduplication with a minimal cost of computation and communication. CACH-Dedup is evaluated using a prototype deployed on windows server environment distributed over four separate machines. It is shown to have duplicate elimination effectiveness in a par with a single node deduplication system, with a minimal communication overhead and an acceptable deduplication throughput.

Read the paper · More papers on PaperTik