Dataset Properties and Degradation of Machine Learning Accuracy with an Anonymized Training Dataset

Wakana Maeda, Toshiya Shimizu, Takeru Fukuoka, Ikuya Morikawa · 2020

With the widespread use of machine learning, we are seeing many use cases where personal data are used for training. There are two privacy threats in such use cases: leakage of training data to a model constructor and inference of personal information from a model by a model user. One countermeasure against these threats is to anonymize training data to preserve privacy. However, it is difficult to find a better anonymization method with less degradation in machine learning performance, such as accuracy, for a given dataset, since anonymization causes inconsistent degradation among datasets. This paper proposes a dataset-property-based method of comparing degradation of machine learning accuracy caused by anonymization. Experiments show that the proposed method can capture the relationship between dataset properties and anonymization methods. They also show the possibility that some dataset properties are related to anonymization and some are not. The proposed method provides a good starting point to help practitioners to choose an anonymization method, that will cause less degradation of machine learning accuracy, for a certain dataset.

Read the paper · More papers on PaperTik