Enhancing Deduplication Efficiency Using Triple Bytes Cutters and Multi Hash Function

International journal of intelligent engineering and systems · 2023

Managing big data backups is challenging due to high volumes of redundant data.Data deduplication is widely used but incurs significant computational and time costs.This paper proposes a hybrid deduplication system that combines file-level and block-level methods to enhance deduplication while reducing costs.File-level deduplication eliminates duplicate files, while block-level deduplication is applied to non-duplicated files using a dynamic list of divisors to enhance deduplication.A multi-hash function generates three hash values for each file or chunk to improve chunking speed and reduce hash collisions.The proposed hybrid system outperforms other state-ofthe-art methods in terms of time, data deduplication ratio, and deduplication gain.Experimental results show reductions of 97.2%, 91.6%, and 82.1% in data size for Dataset 1, Dataset 2, and Dataset 3, respectively, and demonstrate that the proposed multi-hash function is faster and requires less storage than other hash functions.

Read the paper · More papers on PaperTik