Trends Matching as a Dataset Attack Detection Strategy During Machine Learning

Dakalo David Ramalivhana, Colin Chibaya, Kudakwashe Madzima · 2021

Machine learning systems rely on the validity and integrity of the datasets provided to make correct predictions. However, datasets attacks are real. Datasets can be poisoned, spoiled, corrupted, or falsified to mislead machine learning system. Strategies for preventing datasets poisoning are, therefore, apparently needed. Models to defend datasets are scarce. In our context, datasets poisoning is about addition of malicious data into credible datasets during machine learning. We propose a strategy for preventing addition of malicious data into accepted datasets by verifying and matching the trends depicted in credible datasets to the patterns arising from the additional data. Precisely, we investigate the strength of association between the trends depicted in the two datasets before a recommendation to accept or reject the additional data into a machine learning process is made. The age detection datasets, gender classification datasets and uber pickups datasets were used for illustration purposes. These datasets are all freely available online on Kaggle. The age detection dataset was the selected credible dataset. The uber pickups was used as the malicious dataset, while the gender classification dataset was used as the valid additional data. An experimental assessment of the validity of the proposed trends matching strategy indicates that only those datasets with matching trends can be accepted as additional data into credible datasets. Malicious datasets are blocked from poisoning credible datasets. This intervention implicitly supports high quality and accurate machine learning.

Read the paper · More papers on PaperTik