Countermeasure of Polluting Health-Related Dataset for Data Mining
I‐Hsien Liu, Jung-Shian Li, Yen-Chu Peng, Meng-Huan Lee, Chuan-Gang Liu · 2022
Nowadays, machine learning is widely used in a variety of applications, but it still faces many security challenges. Among them, the security of dataset is particularly important, because the data set is the key factor in achieving high correctness for machine learning. Recently, it becomes more difficult for an attacker to directly modify or attack the machine learning models because these models are setup usually in a well-known and well-designed format. However, the attackers can easily manipulate the dataset in various ways. Therefore, we develop countermeasures of polluting o a health-related dataset for data mining, which is robust Data Washing, an algorithm based on denoising autoencoder. It effectively alleviates damages to datasets caused by poisoning attack. We implement several DNN models for different datasets. The proposed Our robust Data Washing algorithm efficiently recovers the poisoning dataset and detect several attacks with a high accuracy rate.