How to Escape a Barnaby: Output Attractors in Large Language Models

Yuki Oyama · Zenodo (CERN European Organization for Nuclear Research) · 2026

Large language models produce strikingly homogeneous creative outputs. We demonstrate that this homogenization is far more severe than previously documented: when asked to invent a dragon for a children's picture book, Google's Gemma 4 (26B and 12B) produces a dragon named "Barnaby" in 100% of independent runs (50/50). Alibaba's Qwen models converge on the same character in 64–80% of runs. A separate meme-invention task yields a golden retriever meme in 48–64% of runs. We term this phenomenon the Barnaby problem: deterministic convergence on specific output patterns despite stochastic sampling, observed across model families, architectures, and tasks. We introduce Barnaby escape, a prompt-only method requiring exactly two LLM calls: one to generate a baseline output that exposes the attractor, and one to regenerate with that output presented as a pattern to avoid. Across four models (9B–35B, MoE and dense), two model families (Google, Alibaba), and two tasks, a single negative example achieves 87.5–100% escape (3 trials × 50 samples per condition). The method requires no additional training, no decoding modification, and no access to model internals. Applied recursively, the method reveals a hierarchical attractor onion: escaping one attractor exposes the next, peelable at a cost of one additional call per layer. Sensitivity analysis confirms that one negative example per layer is optimal. All code, prompts, and experimental data will be released by the time of publication.

Read the paper · More papers on PaperTik