Ignorance is Bliss: Exploring Defenses Against Invariance-Based Attacks on Neural Machine Translation Systems

Akshay Chaturvedi, Abhisek Chakrabarty, Masao Utiyama, Eiichiro Sumita, Utpal Garain · IEEE Transactions on Artificial Intelligence · 2021

This article addresses an invariance-based attack on the transformer, a state-of-the-art neural machine translation (NMT) system. Such attacks make multiple changes to the source sentence with the goal of keeping the predicted translation unchanged. Since thegold translationis not available for the adversarial sentences, tackling invariance-based attacks is a challenging task. We propose two contrasting defense strategies for the same,learn to dealandlearn to ignore. Inlearn to deal, NMT system is trained not to predict the same translation for a clean text and its noisy counterpart, whereas inlearn to ignore, NMT system is trained to output adummy sentencein the target language whenever it encounters a noisy text. The experiments on two language pairs, English–German (en–de) and English–French (en–fr), show thatlearn to dealstrategy reduces the attack success rate from 84.0% to 62.2% for en–de and from 84.6% to 73.8% for en–fr, whereaslearn to ignorestrategy reduces the attack success rate from 84.0% to 27.2% for en–de and from 84.6% to 37.0% for en–fr.

Read the paper · More papers on PaperTik