Adversarial Testing of LLMs Across Multiple Languages

Yulia Kumar, Christopher Paredes, Gang Yang, Jiatao Li, Patricia A. Morreale · 2024

This study builds on prior research in the field of jailbreaking, focusing on the vulnerabilities of state-of-the-art Large Language Models (LLMs) and their associated chatbots. The primary objective is to evaluate these vulnerabilities through a multilingual lens, extending the scope of previous monolingual studies. The main chatbots under examination include ChatGPT (both legacy ChatGPT -3.5 and the latest ChatGPT -40 model), Gemini, Microsoft Copilot, and Perplexity. Researchers conducted multilingual adversarial attacks, facilitating cross-language and cross-model comparisons, and explored different modalities, such as text versus speech inputs. The findings reveal significant weaknesses in these major models, particularly in their susceptibility to adversarial attacks conducted in languages such as Spanish, Russian, and Traditional Chinese. Given the global proliferation and accessibility of these models, it is imperative to rigorously assess the robustness of LLMs against adversarial inputs across multiple languages and modalities.

Read the paper · More papers on PaperTik