Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
Weitao Feng, Lixu Wang, Lv, Peizhuo, Wei, Tianyi, Jie Zhang, Gao, Chongyang, Sinong Zhan, Dong, Wei · arXiv (Cornell University) · 2025
As LLMs become more capable, harmful misuse through fine-tuning becomes increasingly risky. While prior work largely focuses on supervised fine-tuning (SFT), we show that reinforcement learning (RL) enables stronger safety alignment bypasses and more advanced harmful task assistance under matched compute budgets. To address this threat, we propose Token Buncher, the first defense designed specifically against RL-based harmful fine-tuning. This repository accompanies the Token Buncher paper and provides code to launch harmful-RL attacks, apply the defense, and evaluate safety and capability benchmarks.