Energy big data automatic desensitization model based on Spark parallel computing framework

Dongge Zhu, Rui Ma, Yufeng Chai, Bing Cai, Han Liang · 2021 2nd International Conference on Big Data Economy and Information Management (BDEIM) · 2021

With the vigorous development of the electric power industry, the large amount of data generated has also brought tremendous pressure to information security. Data desensitization model is limited by the computing efficiency and memory capacity of a single computing node, which is difficult to meet the desensitization needs of massive data. Based on Spark parallel computing framework, an automatic desensitization model of energy big data is proposed. Depending on the regular dependence between data sets, the elastic distributed data set is established, and the big data parallel processing framework is established by using Spark framework, which decomposes the single node processing task into multi node data blocks. The distributed anonymization algorithm is designed to divide the buffer tuple to complete the anonymization before desensitization. The reserved number of generalized nodes is determined, and a distributed desensitization model based on cooperative computing of multiple computing nodes is established. The experimental results show that the privacy protection strength of this model is 0.0763 and 0.1007 higher than that of the data desensitization model based on centralized anonymization algorithm and regular specific format when the data set size reaches the maximum of 60G. Therefore, the designed model has good data desensitization ability, and can ensure the privacy of data under the condition of large amount of data.

Read the paper · More papers on PaperTik