A Comparative Study on Various ML Models Using Synthetic Data for Privacy Preservation
Md. Ashraf Uddin, Md. Naimul Ahsan, Mrinmoy Das · 2024
In the contemporary world, data holds paramount importance across various sectors, such as healthcare, finance, and government. The confidentiality of information is a critical concern, leading to a surge in research focused on data privacy. Over the past two decades, numerous techniques have been employed to safeguard data, with synthetic data generation emerging as a prominent method. Organizations, including those in the medical, banking, and government sectors, prioritize data privacy. Consequently, researchers have developed techniques for generating synthetic data based on real datasets. This study aims to assess the accuracy of synthetic data in comparison to real datasets using fundamental machine learning models such as Logistic Regression, Random Forest Classifier, Decision Tree, and Support Vector Machine. The investigation involves three authentic datasets related to cancer patients. The datasets are employed to train and test the machine learning models, and performance metrics such as accuracy, precision, and recall are computed for the real data. Simultaneously, synthetic data is generated using the authentic dataset, and the synthetic dataset is trained and tested using the same machine learning models. Performance metrics are then computed for the synthetic dataset. The study concludes with a comprehensive comparison of accuracy, precision, and recall values between synthetic and real data. The comparison is conducted across the four machine learning models to determine which model yields the most favorable results for each dataset.