LERD, a locality enhanced and resemblance based deduplication scheme for large data sets

Panfeng Zhang, Ke Zhou, Hua Wang · 2015

As one kind of storage technology, deduplicaition is widely deployed in all kinds of storage systems. However, the key problems of duplication, such as data throughput and usage of RAM, have not been perfectly addressed. Especially, with the emergence of cloud storage, traditional deduplication methods are not able to adapt to the velocity characteristic of the large data sets. This paper proposes LERD, a temporal locality enhanced resemblance based Duplication scheme, aiming at rapidly querying duplicated data for large scale data sets. LERD takes advantage of data resemblance and temporal locality of data stream to narrow query range, which not only rise throughput, but also decline usage of RAM. Theoretical analysis and experimental results show that LERD's performance is much better than other state-of-the-art schemes.

Read the paper · More papers on PaperTik