Big Data De-duplication Using Classification Scheme based on Histogram of File Stream
Hashem Bedr Jehlol, Loay Edwar George · 2022
Modern technologies generate millions of data every second. Managing and storing this data is very challenging. Data deduplication reduces storage requirements and data redundancy. It is one of the most effective solutions for Big data storage. This paper proposes a new method to accelerate and improve the deduplication technique. It separates the data into many classes before entering the deduplication system based on the Pearson correlation between the histograms of different data extensions. A new method has been developed to divide the file appropriately by generating a list of divisors based on data repeating patterns. In addition, a mathematical algorithm is proposed to construct the hash functions for each chunk that are twice faster as traditional hash functions (MD5 and SHA-1). As a result, the Deduplication Ratio (DR) of the proposed method is about ten times more potent than the Basic Sliding Window (BSW) method and approximately five times more powerful than the Two Thresholds Two Divisors (TTTD) method.