DPMM: Data Privacy and Memory Management in Big Data Server using Hybrid Hashing Method

Manjula GS, T. Meyyappan · 2022 International Conference on Automation, Computing and Renewable Systems (ICACRS) · 2022

Data deduplication (Dedup) is commonly used in the cloud to save bandwidth and storage space by removing duplicate data sets before sending. Data integrity is protected during the Dedup process by encrypting it before it is reused. This paper proposes a DPMM framework for data privacy and memory management in a big data server using a hybrid hashing method. Some programmes use Apache Hadoop because the Hadoop Distributed File System (HDFS) offers a highly dependable static imitation strategy for processing data. The limitations are each file has a different access rate, using the same imitation factor across the board can weaken the performance. Considering this limitations, a research approach has been proposed to utilize predictive examination for progressively imitating the data. So to attain greater efficiency, Dynamic Data Partial Imitation (DDPI) algorithm is implemented, which helps to avoid excessive memory consumption. In the proposed, we have executed the partial imitation technique (DPMM), i.e., the data is fragmented and stored in four Hadoop servers partially to achieve fault tolerance. The proposed method provides the user-requested files even when there is a problem in any servers by this partial imitation technique. Before the imitation process, the files uploaded are checked for duplication to avoid repetition by implementing Secure Hash Algorithm-1 and 2 (SHA1 & SHA2) that apply the hash code of data to check ownership with code verification, thus reducing the excess consumption of storage. Finally, this method is compared with existing methods of execution time for various file sizes. The implemented result demonstrates that this proposed methodology gives much more efficient performance, redundancy and overhead is received.

Read the paper · More papers on PaperTik