SafeBoundary-LLM: Measuring Safety Boundary Stability in Local Open-Weight LLMs Through Single-Turn Baselines and Multi-Turn Escalation
Andreea Alexandra Anghel, Cătălin Anghel, Emilia Pecheanu, Antonio Stefan Balau, Marian Crăciun, Adina Cocu, Cristian Sandu · Computers · 2026
Local open-weight large language models (LLMs) are increasingly used in privacy-sensitive settings, yet isolated prompts may not reveal whether safety boundaries remain stable during conversation. SafeBoundary-LLM evaluated seven local models across 14 sensitive domains, 84 boundary sets, 672 single-turn prompts, and 84 five-turn escalation conversations; the same models were evaluated separately on XSTest and JBB-Behaviors. Evaluators R1 and R2 independently classified all 12,194 responses, with R2 labels used for primary outcomes and unreconciled labels used for reliability analysis. Exact agreement exceeded 93% in each dataset. In SafeBoundary-LLM, 456 out of 7644 responses (5.97%) were confirmed-or-mixed failures. The multi-turn failure rate was 14.69% versus 0.51% for single-turn prompts, yielding a rate ratio of 28.80 (95% CI [21.20, 43.80]; Holm-adjusted p = 0.0006); boundary collapse occurred only at Turns 4–5, and role-play bypass accounted for 299 out of 456 failures. On answer-expected items, over-refusal was 4.34% in XSTest and 17.29% in JBB-Behaviors, whereas unsafe compliance on refusal-expected items was 0.36% and 0.86%, respectively. These findings support an evaluation strategy that includes public single-turn benchmarks, controlled multi-turn escalation, independent human review, and traceable audit records for locally deployed LLMs.