Guardrail Blind Spots on Turkish Prompt Injection: A Same-Set Miss-Rate Comparison of Four Guard Models
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Hu, Qing, Brian B. Fuller, Davide Testuggine, Madian Khabsa · arXiv (Cornell University) · 2023
Self-deposited technical report — not peer-reviewed. On a single fixed set of 248 Turkish prompt-injection attacks, one widely used open guardrail model, jackhhao/jailbreak-classifier, failed to flag 85.48% (212/248) of the attacks, while two other open guards evaluated on the exact same inputs missed almost nothing: fmops/distilbert-prompt-injection missed 0.0% (0/248) and protectai/deberta-v3-base-prompt-injection-v2 missed 1.61% (4/248). An in-distribution baseline detector we trained (altaysec-detector) missed 0.81% (2/248). The headline finding is not that any one guard is "broken," but that the miss rate on Turkish is overwhelmingly a per-model property: on identical inputs the gap between the worst and best guard was over 85 percentage points. We further probe robustness with light obfuscation. Averaged across guards, the miss rate rises from 21.26% on original-form attacks (n=107) to 25.0% when the same attacks are broken up with zero-width spacing (n=21), with casefolding (22.92%, n=60) and leetspeak (21.25%, n=60) in between. We frame these numbers as a disclosed-limit measurement, not a vendor verdict: one of the four models is a jailbreak-specific classifier and our own baseline is in-distribution, so neither is a fair head-to-head winner. The practical message for teams deploying LLMs on Turkish traffic is concrete — you must measure your specific guard on Turkish inputs, because English benchmarks do not transfer. Türkçe özet: Aynı sabit 248 Türkçe prompt-injection saldırısı üzerinde, yaygın kullanılan açık bir koruma modeli jackhhao/jailbreak-classifier saldırıların %85,48'ini (212/248) yakalayamadı; oysa aynı girdiler üzerinde fmops/distilbert-prompt-injection %0,0 (0/248), protectai/deberta-v3-base-prompt-injection-v2 ise %1,61 (4/248) kaçırdı. Eğittiğimiz taban çizgisi altaysec-detector %0,81 (2/248) kaçırdı. Ana bulgu, herhangi bir korumanın "bozuk" olması değil, kaçırma oranının büyük ölçüde modele özgü olmasıdır: aynı girdilerde en kötü ile en iyi arasındaki fark 85 puanı aştı. Hafif gizleme ile ortalama kaçırma orijinalde %21,26'dan sıfır-genişlikli boşlukta %25,0'a çıkıyor. Kısıtlar açıkça belirtildi: modellerden biri jailbreak'e özgü, kendi taban çizgimiz ise dağılım içidir. Pratik ders nettir — İngilizce iyi olan bir koruma otomatik olarak Türkçe iyi değildir; kendi korumanızı Türkçe girdilerle ölçün. Open data & code: GitHub (code & data, open) · Hugging Face dataset · companion article: altaysec.com.tr. Author: Fevzi Ege Yurtsevenler, AltaySec (Türkiye). License: CC BY 4.0.