AltayDuel: A Turkish-First Arena and Open Dataset for Multi-Turn LLM Prompt-Injection Red-Teaming
Fevzi Ege Yurtsevenler · Zenodo (CERN European Organization for Nuclear Research) · 2026
Prompt injection is the top-ranked risk in the OWASP Top 10 for Large Language Model (LLM) Applications, yet the red-team corpora that drive defenses are overwhelmingly English. Languages with different morphology and cultural framing are under-represented. We present AltayDuel, an agent-vs-agent self-play arena in which an attacker ("red") LLM attempts, over multi-turn dialogue (1–8 rounds, mean 4.6), to make a defender ("blue") LLM violate its system instructions, adjudicated by a deterministic judge. From the arena we release two openly licensed (CC-BY-4.0), Turkish-first datasets: (i) 2,594 cleaned multi-turn duel transcripts containing 439 judge-confirmed successful attacks, and (ii) a curated set of 300 categorized Turkish prompt-injection payloads, balanced at 25 across 12 attack families and mapped to the OWASP LLM Top 10 and MITRE ATLAS. Analysing the transcripts, the overall attack-success rate is 16.9% (439/2,594); success is comparable across languages — 16.0% on Turkish scenarios and 18.2% on English — and successful attacks manifest predominantly as capitulation under pressure (51.3%) and partial leakage (41.7%) rather than verbatim secret disclosure (7.1%). We contribute a taxonomy of Turkish-specific evasion vectors (morphological manipulation, code-switching, culturally-framed authority) as a catalogued attack surface for locale-aware guardrails, and release the datasets as reproducible Turkish-language safety benchmarks.