Comprehensive Analysis of Content Defined De-Duplication Approaches for Big Data Storage

C. Vijesh Joe, S. Shinly Swarna Sugi · 2022 Sixth International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC) · 2022

In day-to-day life, huge amounts of data are generated and storage of that data becomes a difficult task. Archive and backups are the storage medium used to store the data. As the data in the backup and archive are redundant, it needs more storage space. Data deduplication methodology is utilized to shrink the storage space. De-duplication is the technique to remove redundant data and increase the storage space in the storage medium. To store the data in the storage medium without redundancy various Chunking algorithms are used. These algorithms can improve the availability of storage space in the Warehouse. This paper presents the performance evaluation of Two Threshold Chunking (TTC) and Rapid Asymmetric Maximum (RM in Content Defined Chunking Algorithm. Also, the performance of those algorithms is measured by several parameters like De-duplication Elimination Ratio (DER), Throughput and Processing time, etc. Based on the analysis, the processing time of TTD is less and throughput is high compared to RAM, which is discussed in detail with an experimental analysis in this paper.

Read the paper · More papers on PaperTik