safety-tuningEMNLP 2026 Main ✓
Minji Kim, Hyounghun Kim · arXiv:2609.04714 · Sep 4, 2026
Decomposing safety responses into boilerplate vs. rationale reveals which component drives over-refusal.
Finds that safety-tuning over-refusal is largely driven by the boilerplate refusal statement (not the rationale) because it trains the model to pattern-match on superficial cues. Training only on rationales (removing the "I cannot..." boilerplate) reduces false refusal rates while maintaining comparable genuine safety performance — a cheap, high-leverage fix for aligned models that are overly conservative on benign inputs.
Uses normalized activation-patching recovery scores (causal signals from mechanistic interpretability) to select which attention heads and MLP blocks receive LoRA adapters, rather than using standard heuristics (magnitude, activation norm, gradient/Fisher). Evaluated on Ministral-8B/NF4 on Swahili span-JSON information extraction, MechSparse outperforms all heuristic baselines at matched parameter budgets — a concrete applied demonstration that mechanistic causal signals improve practical fine-tuning efficiency.