A Study on the Effect of Data Generalization on Deep Learning Performance

Sung-Bong Jang, Young Woong Ko · Asia-pacific Journal of Convergent Research Interchange · 2022

k-anonymization is a technology that changes original data and distributes it to third parties to prevent hackers from stealing personal information.Data generalization is widely used in kanonymization where the data values are replaced by high-level ones.Although this method has been proven to be very effective in protecting personal information, when generalized data is used as machine learning data such as deep learning, there is a problem that the learning performance is degraded.This study tried to find out how much the value distortion caused by data generalization affects deep learning performance.In the experiment, a deep learning model was built using TensorFlow, the same model was trained using a non-generalized dataset and a generalized dataset, and the learning performance of each was measured.The Korea's temperature collected from January 1, 2010 to June 2022 was used as training dataset.In addition, in order to generate distorted data, the values were changed into high-level generalized ones by force.Root Mean Square Error (RMSE) was used as the performance metric value.The experiment results show that when the degree of generalization was small, there was a performance degradation of 0.90%, when the degree of generalization was medium, a performance degradation of 3.15%, and when the degree of generalization was severe, a performance degradation of 6.57% occurred.

Read the paper · More papers on PaperTik