Acoustic and Visual Approaches to Adversarial Text Generation for Google Perspective

Stephan Brown, Petar Milkov, Sameep Patel, Yi Zen Looi, Ziqian Dong, Huanying Gu, N. Sertac Artan, Edwin Jain · 2019

Google's Perspective API was introduced to help detect and classify toxic comments in online platforms. Adversarial machine learning attacks can decrease the effectiveness of Perspective in identifying toxic comments. We have shown in our previous study that by applying a semantic-based attack to a surrogate model trained with just 10,000 queries could produce adversarial examples which evade Perspective 25% of the time. In this paper, we propose two new approaches to generate adversarial text to evade Google's Perspective, one based on acoustic similarity and the other based on visual similarity. We tested the success rate of obfuscation in Google Perspective using the adversarial texts generated through the proposed approaches and showed that Google Perspective misclassified the generated texts 33% and 72.5% of the time for the visual-based and acoustic-based approaches, respectively. The study aims to broaden the understanding of adversarial text generation and to improve the robustness for online toxic comment detection for a safe online community.

Read the paper · More papers on PaperTik