Exploiting the Data Redundancy Locality to Improve the Performance of Deduplication-Based Storage Systems

Suzhen Wu, Xiao Chen, Bo Mao · 2016

The chunk-lookup disk bottleneck and the read amplification problems are two great challenges for deduplication-based storage systems and restrict the applicability of data deduplication for large-scale data volumes. Previous studies and our experimental evaluations have shown that the amount of redundant data shared among different types of applications is negligible. Based on the observations, we propose AA-Plus which effectively groups the hash index of the same application together and divides the whole hash index into different groups based on the application types. Moreover, it groups the data chunks of the same application together on the disks. The extensive trace-driven experiments conducted on our lightweight prototype implementation of AA-Plus show that compared with AA-Dedupe, AA-Plus significantly speeds up the write throughput by a factor of up to 6.9 and with an average of 3.1, and speeds up the read throughput by a factor of up to 3.3 and with an average of 1.9.

Read the paper · More papers on PaperTik