The Influence of Synonymous Data Transformation on the Performance of Machine Learning Models

Н.Г. Бабак, Leonid Belorybkin, Shamil Otsokov, Mikhail Poletaev, Aleksey Terenin, Anastasiya Shabrova · Vestnik MEI · 2024

A problem lying in the fact that in some machine learning applications, personal data become unsuitable for use after they have been anonymized. The purpose of the analysis is to study the possibilities of synonymous anonymization---a processing that has to be carried out to comply with the requirements of the federal law on personal data protection---for preserving the quality of machine learning models. The effect of training deep learning models on anonymized data using classical anonymization and synonymous transformation methods is analyzed, and the performance metrics of these models are compared with similar models trained on personal data. It has been found that the use of classical anonymization methods resulted in that the performance of machine learning models became degraded by 33% on the average, while models trained on synonymously anonymized data showed a quality commensurable with that of models trained on personal data. Synonymous data transformation has been proposed as an effective data anonymization approach for machine learning, which makes these data more available for analysis and research purposes without compromising the performance and reduces the risks associated with the processing and transfer of personal data.

Read the paper · More papers on PaperTik