Usability of a synthetically generated dataset for decision support
Oliver Lohaj, Ján Paralič, Daria Kushnir, Jakub Ivan Vanko · 2024
This article deals mainly with the issue of usability of a synthetically generated data set. The work focuses on providing a brief overview of synthesizing medical data and presenting suitable methods for evaluating their quality. The work is divided into a theoretical and a practical part, where the theoretical part contains basic concepts, methods and metrics, and the practical part deals with the generation of synthetic data and their qualitative evaluation. One of the most suitable methods for generating synthetic data is the GAN (Generative Adversarial Network) method. We used the CTGAN model for generating synthetic data set from electronic health records. We have evaluated the generated synthetic data with various metrics, like accuracy of models trained on synthetic data, similarity with original data and security of synthetic data. The accuracy of models trained on synthetic data decreased a bit, e.g., from 0.9444 to 0.8704 with Random Forest classifier, in comparison with the original data. The similarity value reached 0.9044, but all the rows in synthetic data were new and none match the original data set, which is important in case of anonymization patients’ data. We also received a resulting privacy value for two model situations of 0.9988 and 0.6507 resp., which show that data has some level of privacy, but it can be improved. The CTGAN method was deemed usable.