Using file-aware deduplication to improve capacity in storage systems

Paul Bartus, Emmanuel Arzuaga · 2017

Storage systems contain redundant copies of data such as identical files or within sub-file regions. Using deduplication technology, we can take advantage of this redundancy and reduce the space needed to store files in the file systems. Scalable, highly reliable distributed systems supporting data deduplication have recently become popular for storing backup and archival data. There is potential for this technology to be adapted to primary storage. The purpose of our research is to create a file-type-aware deduplication system to improve capacity in storage systems. We have initiated our study with focus on the relation between the amount of duplicate content that we can extract among different files to understand the relation between duplicate content and file type. We are particularly interested in the case in which once a file type is given, we can provide the appropriate parameters (chunk size, id) to apply deduplication on it Given the differences between file structures, this mechanism can be used to create a file-aware deduplication system.

Read the paper · More papers on PaperTik