SentMod: Hidden Backdoor Attack on Unstructured Textual Data

Saquib Irtiza, Latifur R. Khan, Kevin W. Hamlen · 2022

For a very long time, machine learning techniques such as deep learning have been considered to be very reliable and accurate for a wide variety of tasks including security sensitive applications. But recently, it was found that such models were vulnerable to adversarial attacks. It caused these models to generate wrong results when perturbed data was passed through it. One such attack is the backdoor attack where a backdoor is formed in the model when it is trained on a dataset containing malicious data. The malicious data is formed by adding small predefined perturbations to clean instances and changing its label to the target class. The backdoor is usually hidden and is only activated when the model receives a test data containing the same perturbation that was added to the training dataset. In natural language processing, the perturbations are usually a sequence of trigger words that are added to a predefined position in the sentence. But, these perturbations are often very obvious and can be detected by manual inspection since the context of the sentence does not match with its label. So, we have come up with a novel backdoor attack strategy that can generate perturbed data in a way such that its context matches with its label and the perturbation is hidden better. Our algorithm called SentMod, perturbs only 2% of the training data to obtain an attack success ratio of 97%. We have tested with multiple deep learning models to show that our approach is highly effective and is also very difficult to be detected using most common defense methods.

Read the paper · More papers on PaperTik