02
refusal geometry mech-interp preprint · Aug 25

Refusal Geometry Reflects Refusal Training: Diverse Refusal Prefixes Can Raise Stable Rank and Weaken Refusal Vector Ablation Attacks

Abliteration works by extracting a single refusal direction and orthogonally projecting it out. This paper shows that the underlying vulnerability is a training artifact: when all refusals start the same way, refusal mass concentrates in one direction — which is exactly what abliteration exploits. Diverse refusal prefixes spread refusal across many directions, making it geometrically much harder to ablate.

Refusal prefix diversity → Stable rank → stable rank ablation ASR ↑ diversity → ↑ stable rank → ↓ abliteration success
Figure 2: Higher refusal prefix diversity during training raises the stable rank of refusal representations (blue), which simultaneously weakens refusal vector ablation attacks (red dashed). Low-diversity training produces a low-rank refusal subspace — easily ablated by orthogonal projection.

Abliteration extracts the refusal direction from a small set of contrastive prompts and removes it via weight projection. The stable rank of the refusal subspace determines how many orthogonal directions would need to be removed to fully suppress refusal. This paper demonstrates empirically that training with diverse refusal prefixes — varied sentence starters for refusal outputs — distributes refusal mass across more directions, raising stable rank and significantly increasing the number of projections an attacker must perform. The result is a training-time hardening prescription against a family of white-box alignment attacks that requires no change to the post-training stack other than curating diverse refusal prefix data.

03
refusal direction weight editing preprint · Aug 2026

Abliteration Mitigation via Refusal Aliases (AMRA)

Abliteration works because the refusal direction is reliably extractable via contrastive prompting. AMRA attacks that extractability directly: it replaces the activations that encode the refusal direction with random aliases in the writer matrices, making the direction attackers extract meaningless while preserving all model behavior through corrected reader matrices.

Writer W refusal direction rank-k update ↳ alias injection W′ (aliased) random alias ≠ refusal reader correction → behavior preserved attacker extracts alias → ablates nothing
Figure 3: AMRA applies a rank-k update to writer matrices, replacing refusal-encoding activations with random aliases. Reader matrices are corrected to preserve model behavior. When attackers run contrastive extraction, they obtain the alias — not the real refusal direction — so ablation fails. Llama-3-8B: +2.16 refusal score post-abliteration, <0.5pp MMLU cost.

AMRA intervenes at the weight level before deployment. Rank-k updates to residual stream writer matrices replace the activation patterns that ordinarily encode the refusal direction with random alias vectors, while downstream reader matrices are corrected via a matching update to restore original behavior (for all content other than the alias). An attacker running contrastive extraction obtains the alias vector; ablating it removes nothing safety-relevant, leaving refusal intact. On Llama-3-8B, AMRA improves post-abliteration refusal scores by 2.16 points over the undefended baseline with less than 0.5 percentage points of MMLU degradation, confirming that alias injection is achievable with minimal capability cost.

Items 4 – 10  ·  Also notable
04
AI security mech-interp preprint · Aug 14

Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

per-neuron hyp. tests FDR control safety neurons certified S⊆N clamp → harmful- conditional mean ≤2.0% ASR · 0.5–5.3% Δutil Step 1 Step 2 Step 3 (inference / weight edit)
Figure 4: Tripwire pipeline — FDR-controlled neuron hypothesis tests → certified safety neuron set → inference-time clamp (or offline bias-patch). ASR ≤ 2.0% across 4 models × 4 attacks.

Training-free defense: per-neuron hypothesis tests under FDR control identify safety-specific neurons; a clamp holds them at their harmful-conditional mean activations at inference time (equivalently implemented as an offline bias-patch weight edit). Reduces average attack success rate to ≤2.0% across four safety-aligned LLMs and four attack types, with only 0.5–5.3% MT-Bench utility drop — smallest utility cost among all compared defenses.

05
ACL 2026 mech-interp over-refusal

SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering

steer layer depth → refusal prob harmful → refusal ✓ safe → over-refused ✗ after steering ✓
Figure 5: Task-specific "constellation" trajectories across layers — harmful prompts (red) and over-refused safe prompts (orange dashed) follow different paths. SafeConstellations applies task-conditioned steering (blue) to redirect over-refused representations toward non-refusal trajectories. ACL 2026 Long Paper #2056.

Mechanistic analysis reveals that LLM representations follow task-specific "constellation" trajectories across layers with distinct refusal vs. non-refusal paths. SafeConstellations tracks these patterns and applies dynamic-layer task-conditioned steering to redirect over-refused representations, reducing over-refusals by up to 73% while causing less disruption to the harmful-refusal mechanism than global ablation methods. Peer-reviewed at ACL 2026.

09
agent safety preprint · Aug 20

SafeBranch: Branch-Pair Safety Alignment for Embodied Agents

unsafe rollout safety-critical step unsafe action safe alternative DPO branch-pair
Figure 9: SafeBranch rolls back unsafe rollouts to the safety-critical step, obtains a safe alternative, and pairs the two branches as a DPO preference example differing only at the critical action. Training improves embodied agent safety while preserving task completion.

VLM-based embodied agents often violate safety constraints at isolated critical steps within otherwise benign trajectories. SafeBranch rolls back each unsafe rollout to the safety-critical step, queries the actor for a safe alternative action, then constructs a DPO preference pair contrasting unsafe vs. safe under identical trajectory context. This provides targeted safety alignment without human annotation or trajectory-level labeling.

10
alignment eval preprint · Aug 23

Who Pays More for Safety? Measuring the Disparate Cost of Safety Alignment across Languages

EN 0.8 ES 1.2 FR 1.5 ZH 1.8 AR 2.1 HI 2.3 Safety Cost → Non-English users bear systematically higher Safety Cost
Figure 10: "Safety Cost" (SC) — utility loss from safety alignment — by language group. Non-English users (ES, FR, ZH, AR, HI) consistently bear higher SC than English users, revealing a systematic inequity in current safety alignment practice.

The "Safety Cost" (SC) metric measures the utility loss imposed by safety alignment across language groups. Applied across multiple aligned models, SC reveals a systematic inequity: non-English users consistently bear a higher Safety Cost than English users, indicating that current safety training practices implicitly trade off non-English utility for English-centric safety calibration. Provides both a measurement protocol and a diagnostic for multilingual alignment failures.

4 entries removed on 2026-09-10 as repeats of earlier reports: 2606.19168 (first covered 2026-08-28), 2608.08383 (first covered 2026-08-28), 2608.26008 (first covered 2026-08-28), 2608.02674 (first covered 2026-08-11).

← all Research Radar issues · gussand · source