Data Anonymization Using Imbalanced Data for Deep Learning with Uppersampling and Undersampling
Ayahiko Niimi · International Journal of Intelligent Computing Research · 2019
Most real data are imbalanced and thus present difficulties as a subject of classification problems.Some research has been conducted on the topic of imbalanced data.However, managing such data in deep learning has not been considered.In this study, we propose privacy-protection data mining using imbalanced data through deep learning.We discuss existing privacy-protection data mining, study its features, and examine a deep learning based anonymizing tool.Experiments using various anonymization tools confirm that deep learning does not reduce data accuracy by making that data anonymous.Under and uppersampling are both used for imbalanced data, but the tendencies of these approaches to reduce data accuracy because of anonymization have not been confirmed. Input DataWe first consider the "anonymization" and "randomization" of representative methods of input-