Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
Refusal-Escape Directions (RED) are local continuous perturbation directions near harmful inputs that shift model behavior from refusal to compliance while preserving harmful-semantics interpretation — giving a mechanistic, operator-level account of why jailbreaks structurally work. The framework proves that fine-tuning-based alignment cannot eliminate RED without capability loss, establishing a fundamental safety-utility trade-off that output-level defenses cannot resolve.