Abliteration works by extracting a single refusal direction and orthogonally projecting it out. This paper shows that the underlying vulnerability is a training artifact: when all refusals start the same way, refusal mass concentrates in one direction — which is exactly what abliteration exploits. Diverse refusal prefixes spread refusal across many directions, making it geometrically much harder to ablate.
Figure 2: Higher refusal prefix diversity during training raises the stable rank of refusal representations (blue), which simultaneously weakens refusal vector ablation attacks (red dashed). Low-diversity training produces a low-rank refusal subspace — easily ablated by orthogonal projection.
Abliteration extracts the refusal direction from a small set of contrastive prompts and removes it via weight projection. The stable rank of the refusal subspace determines how many orthogonal directions would need to be removed to fully suppress refusal. This paper demonstrates empirically that training with diverse refusal prefixes — varied sentence starters for refusal outputs — distributes refusal mass across more directions, raising stable rank and significantly increasing the number of projections an attacker must perform. The result is a training-time hardening prescription against a family of white-box alignment attacks that requires no change to the post-training stack other than curating diverse refusal prefix data.
03
refusal directionweight editingpreprint · Aug 2026
Abliteration works because the refusal direction is reliably extractable via contrastive prompting. AMRA attacks that extractability directly: it replaces the activations that encode the refusal direction with random aliases in the writer matrices, making the direction attackers extract meaningless while preserving all model behavior through corrected reader matrices.
Figure 3: AMRA applies a rank-k update to writer matrices, replacing refusal-encoding activations with random aliases. Reader matrices are corrected to preserve model behavior. When attackers run contrastive extraction, they obtain the alias — not the real refusal direction — so ablation fails. Llama-3-8B: +2.16 refusal score post-abliteration, <0.5pp MMLU cost.
AMRA intervenes at the weight level before deployment. Rank-k updates to residual stream writer matrices replace the activation patterns that ordinarily encode the refusal direction with random alias vectors, while downstream reader matrices are corrected via a matching update to restore original behavior (for all content other than the alias). An attacker running contrastive extraction obtains the alias vector; ablating it removes nothing safety-relevant, leaving refusal intact. On Llama-3-8B, AMRA improves post-abliteration refusal scores by 2.16 points over the undefended baseline with less than 0.5 percentage points of MMLU degradation, confirming that alias injection is achievable with minimal capability cost.
Figure 4: Tripwire pipeline — FDR-controlled neuron hypothesis tests → certified safety neuron set → inference-time clamp (or offline bias-patch). ASR ≤ 2.0% across 4 models × 4 attacks.
Training-free defense: per-neuron hypothesis tests under FDR control identify safety-specific neurons; a clamp holds them at their harmful-conditional mean activations at inference time (equivalently implemented as an offline bias-patch weight edit). Reduces average attack success rate to ≤2.0% across four safety-aligned LLMs and four attack types, with only 0.5–5.3% MT-Bench utility drop — smallest utility cost among all compared defenses.
Figure 5: Task-specific "constellation" trajectories across layers — harmful prompts (red) and over-refused safe prompts (orange dashed) follow different paths. SafeConstellations applies task-conditioned steering (blue) to redirect over-refused representations toward non-refusal trajectories. ACL 2026 Long Paper #2056.
Mechanistic analysis reveals that LLM representations follow task-specific "constellation" trajectories across layers with distinct refusal vs. non-refusal paths. SafeConstellations tracks these patterns and applies dynamic-layer task-conditioned steering to redirect over-refused representations, reducing over-refusals by up to 73% while causing less disruption to the harmful-refusal mechanism than global ablation methods. Peer-reviewed at ACL 2026.
Figure 9: SafeBranch rolls back unsafe rollouts to the safety-critical step, obtains a safe alternative, and pairs the two branches as a DPO preference example differing only at the critical action. Training improves embodied agent safety while preserving task completion.
VLM-based embodied agents often violate safety constraints at isolated critical steps within otherwise benign trajectories. SafeBranch rolls back each unsafe rollout to the safety-critical step, queries the actor for a safe alternative action, then constructs a DPO preference pair contrasting unsafe vs. safe under identical trajectory context. This provides targeted safety alignment without human annotation or trajectory-level labeling.
Figure 10: "Safety Cost" (SC) — utility loss from safety alignment — by language group. Non-English users (ES, FR, ZH, AR, HI) consistently bear higher SC than English users, revealing a systematic inequity in current safety alignment practice.
The "Safety Cost" (SC) metric measures the utility loss imposed by safety alignment across language groups. Applied across multiple aligned models, SC reveals a systematic inequity: non-English users consistently bear a higher Safety Cost than English users, indicating that current safety training practices implicitly trade off non-English utility for English-centric safety calibration. Provides both a measurement protocol and a diagnostic for multilingual alignment failures.
4 entries removed on 2026-09-10 as repeats of earlier reports: 2606.19168 (first covered 2026-08-28), 2608.08383 (first covered 2026-08-28), 2608.26008 (first covered 2026-08-28), 2608.02674 (first covered 2026-08-11).