Lei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song, Tsung-Yi Ho, Pin-Yu Chen, Yaoqing Yang · ACL 2026 main · arXiv 2506.05346
Research on guardrail collapse has obsessed over fine-tuning methods — SFT vs. RLHF vs. DPO — while ignoring the upstream alignment dataset itself. This ACL long paper shows that dataset similarity is the missing variable: if your safety alignment data looks like the fine-tuning data, your guardrails fail regardless of how carefully you fine-tune.
Figure 3: Dataset similarity vs. post-fine-tuning harmfulness — high similarity predicts guardrail collapse; low-similarity alignment data reduces harmfulness by up to 10.33%.
The authors measure representation similarity (CKA-style) between models' upstream safety alignment datasets and downstream fine-tuning task datasets, then correlate this metric with post-fine-tuning harmfulness scores on jailbreak benchmarks. High alignment-task similarity significantly weakens guardrails; choosing alignment data with low similarity to foreseeable fine-tuning domains reduces harmfulness by up to 10.33% while maintaining utility. The finding reframes guardrail collapse as an alignment-data design problem rather than purely a fine-tuning-method problem: providers should diversify their alignment datasets to reduce overlap with likely downstream task distributions.