E-DAID: An Efficient Distributed Architecture for In-Line Data De-duplication

Seetendra Singh Sengar, Manoj Kumar Mishra · 2012

As data have been growing rapidly in data centers, data de-duplication, a form of compression, has become an important need of most commercial and research backup systems. Currently, data de-duplication storage systems continuously facing challenges in providing the required throughputs and capacities necessary to move backup data within backup and recovery window times. In this paper, we are presenting a distributed architecture for in-line data de-duplication with one node designated as server and multiple storage nodes. In proposed architecture, we used an Intelligent Storage Balancing Strategy to distribute the data among the storage nodes to improve the de-duplication efficiency. All the nodes, including the server can do block level de-duplication in parallel. Proposed architecture can de-duplicate with high throughput, support de-duplication ratio comparable to that of a single system. And in the last section of this paper, we are introducing a technique called Sampled Hashing for improving the scalability of the architecture.

Read the paper · More papers on PaperTik