Optimized cloud storage capacity using data hashes with genetically modified SHA3 algorithm

Manpreet Kaur, Anurag Jain, Amit Verma · 2017

Cloud computing is a technology that provides an adaptable, on-demand self-service, speedy and a universal access to various resources and computing properties. A procedure of alleviating need for storage by elimination of redundancy, preserving only a distinct instance of information is known as deduplication. The process that is used in de-duplication involves deleting the duplicate copies of same information, leaving behind a single copy, for storing in a storage medium. The analysis of information is performed to find a duplicate pattern of bytes in order to make sure that only a unique instance of information is a file itself. In case, duplicates are found, they are replaced with references that points to a unique instance of a stored file. This technique of data de-duplication is mostly used for minimizing the storage space and saving bandwidth on the server in cloud computing. However, De-duplication of data, focuses on sections that are static like backup and storage systems, are not completely appropriate for cloud storage system due to the dynamic nature of the data as procedure for storing the data in the cloud keeps on changing spontaneously when similar datasets are updated or accessed at a same time by multiple number of users. In this paper authors have discussed about the issue of data Deduplication authorization. In order to ensure the privacy of data that is highly sensitive during De-duplication, a hashing technique based on SHA3 and concept of Genetic Algorithm has been proposed to find duplicity in the data before storage.

Read the paper · More papers on PaperTik