Improving storage capacity by distributed exact deduplication systems

Cristian Barca, Dan Claudiu Barca, Constantin Mara, Petre Anghelescu, Bogdan Gavriloaia, Radu Vizireanu, Răzvan Crăciunescu, Octavian Fratu · 2015

The topic of data deduplication has received lately a lot of attention for its storage reduction functionality. Data deduplication essentially refers to the elimination of redundant data, leaving only one copy of the data to be stored, and is meant to reduce the pain regarding the exponential data growth in backup or archiving centers. Most existing state-of-the-art deduplication systems rely on approximate deduplication in order to achieve high-performance. Unfortunately, these studies are usually conducted and tested on single-host systems. Although their authors claim that the design can be easily applied on multi-node systems, we have not seen yet an extension that enacts that - they lack of trust. Thus, in a world where data deduplication storage systems are continuously struggling in providing the required throughput and disk capacities necessary to store and retrieve data within reasonable times, we are handled the task to design a distributed deduplication systems that will achieve efficiency, scalability and throughput at a petascale capacity level. In this paper we present a proof-of-concept design that one can use to implement such a system: A Distributed Exact Deduplication System, which we believe it will cross the boundaries towards a new generation of backup and archiving systems.

Read the paper · More papers on PaperTik