📡 Research Radar · Weekly · 2026-W36

Refusal Under Pressure

Window: Aug 24–30, 2026 · extended sweep Aug 8–24
Sources: arXiv cs.CL/cs.LG/cs.CR/cs.AI · COLM 2026 · ICML 2026 Mech Interp Workshop · OpenReview
2 peer-reviewed 9 preprints 0 forum/blog 11 total

Theme of the week

The week's dominant signal is the geometry of the refusal mechanism under adversarial pressure. Three independent papers — Abliteration Mitigation via Refusal Aliases, Decided Upstream Written Late, and Tripwire — simultaneously attack the same question from defense, circuit-tracing, and statistical-certification angles, revealing that the refusal signal is both architecturally fragmented (harm is detected mid-network in a language-invariant direction, but refusal is written late by a localised circuit) and geometrically obfuscatable (the refusal direction is easily extracted and projected out unless actively aliased). On the security side, two Aug 26 submissions bring automated red-teaming and adaptive defense into production range: NeuronFuzz eliminates response-generation from the fuzzing loop by using safety-neuron activations as continuous prefill-only feedback, while a self-evolving defense abstracts attack patterns into method-level rules that generalise across attack families. Text diffusion LMs remain a live mechanistic attack surface: safety neurons in DLLMs are sparse, localised, and transferable from autoregressive predecessors, with self-pruning raising ASR from 2.6% → 73.8% on LLaDA. An ICML 2026 Mech Interp Workshop result caps the week by decomposing eval-awareness into capabilities-framing vs. safety-framing, showing the former predicts compliance with a +24–46 pp gap — with direct implications for red-teaming and monitoring pipeline design.

01

Abliteration Mitigation via Refusal Aliases

arXiv preprint · August 2026

safety refusal mechanism preprint

Abliteration — removing the refusal direction via projection — has become a commodity attack requiring only a handful of contrastive prompts. AMRA is the first defense that addresses the root cause: the extractability of the refusal direction itself.

Undefended refusal dir r̂ contrastive extraction → project r̂ out refusal gone AMRA alias r̃ (writer) true r̂ (reader) extraction gets r̃, not r̂ → project r̃ out refusal intact
AMRA replaces refusal-inducing activations with random aliases in writer matrices; reader matrices corrected to preserve behavior. Contrastive extraction finds the alias, not the true direction. Llama-3-8B: +2.16 post-abliteration refusal score, <0.5 pp MMLU cost.

Abliteration works by extracting the refusal direction r̂ from contrastive harmful/benign prompt pairs, then projecting weight matrices orthogonal to r̂. AMRA applies rank-k updates to residual-stream writer matrices to replace the true refusal-inducing activation with a random alias, then corrects downstream reader matrices to preserve the original model behavior. The net effect: the contrastive extraction procedure finds the alias (which no longer encodes refusal) rather than the true direction, so projection-based removal has no effect on actual refusal behavior. On Llama-3-8B, AMRA improves post-abliteration refusal scores by 2.16 points over the undefended baseline with less than 0.5 percentage points of MMLU degradation.

02

Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE

Ramakrishna P. Kompella, Aadit Mahajan · COLM 2026 Actionable Interpretability Workshop (peer-reviewed)

mech-interp safety peer-reviewed

Safety gaps in multilingual models look like failures to detect harm. They are not. Harm is detected in a language-invariant direction mid-network — but refusal is written by a separate, late, localised circuit whose architecture makes cross-lingual transfer expensive in one direction and cheap in the other.

Detection–Writing Dissociation in Sarvam-30B MoE L0 L11 (mid) L-last harm direction (EN–Indic cos ≈ 0.9): detected mid-network refusal write (late) orthogonal damping opposer: cheap ✓ amplifying writer: cost wall ✗
Harm detection is language-invariant at L11 (cosine ≈ 0.9, English vs. Indic). The refusal write is late, localised to a MoE writer + attention-opposer circuit, and geometrically orthogonal to detection. Damping the opposer is cheap; amplifying the writer hits a cost wall with context length.

Evaluated on Sarvam-30B, an Indic-multilingual mixture-of-experts model. Harm is encoded as a direction nearly language-invariant in mid-network (English–Indic cosine ≈ 0.9 at layer 11); steering this direction causally controls refusal output. However, this detection direction is geometrically orthogonal to the direction that actually writes the refusal token — the write is late in the network and assembled across generation. Circuit attribution localises the write to a specific MoE writer head held in check by an attention opposer. Key interventional asymmetry: damping the opposer is cheap and reliably produces refusal; amplifying the writer hits a cost wall (required intervention strength grows with context length). Cross-lingual safety gaps persist not because harm isn't detected, but because the write signal doesn't transfer across languages — explaining why naive fine-tuning on English refusal examples fails to close the gap.

03

Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

arXiv preprint · August 14, 2026

mech-interp jailbreak defense preprint

Every existing jailbreak defense either perturbs all requests (utility cost) or requires a separate training run (deployment cost). Tripwire identifies a certified compact set of safety neurons via hypothesis testing and patches only those neurons — only when needed.

Tripwire: FDR-Certified Neuron Identification → Targeted Patch per-neuron hypothesis tests (BH-FDR) compact safety neuron set S alarm fires → patch S only → refusal reinstated benign input: no alarm → zero cost · attack input: alarm → S patched
Per-neuron hypothesis tests under Benjamini-Hochberg FDR control identify a compact safety-neuron set S. At inference time the patch fires only on alarmed inputs — zero cost on benign. Tested on 4 models × 4 attack types (GCG, AmpleGCG, AutoDAN, Jailbreak-R1).

Tripwire identifies safety neurons using per-neuron hypothesis tests comparing activations on harmful vs. benign inputs under Benjamini-Hochberg FDR control, with utility-specific criteria excluding neurons that also fire on benign requests. At inference time, when a safety alarm fires (computed from safety-neuron activations during prefill), Tripwire patches only the identified neurons to reinstate their refusal-inducing activation pattern — all other neurons are unmodified, and benign inputs incur zero intervention cost. Evaluated on Llama-2-7B, Llama-3.1-8B, Qwen2.5-7B, and Qwen2.5-32B against GCG, AmpleGCG, AutoDAN, and Jailbreak-R1, compared against RepE, TraceRouter, LED, and DELMAN baselines. Tripwire matches or exceeds state-of-the-art safety defense with lower MT-Bench degradation than weight-editing methods.

04

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

Tongyan Hu, Bryan Hooi (National University of Singapore) · arXiv preprint · August 26, 2026

AI security jailbreak defense preprint

Static defenses fail under adaptive attackers because new attack wrappers require manual updates. This framework instead learns from each failure: it abstracts the structural pattern of a successful attack into a method-level rule that generalises across the entire attack family — no parameter updates required.

Self-Evolving Defense: Cross-Interaction Rule Memory attack succeeds abstract → method-level (structural wrapper, not topic) persistent rule memory retrieve on new input one rule → whole attack family · no param updates · open-weight + black-box
On attack success, the framework abstracts the structural wrapper into a method-level rule stored in persistent cross-interaction memory. Retrieval via semantic similarity at inference time generalises across the entire attack family. Operates through external memory and prompting only — no parameter updates; compatible with black-box APIs.

Because rules are method-level (capturing the structural wrapper — role-play, code-transformation, multi-step indirection — rather than the harmful topic), one induced rule generalises across an entire attack family. The mechanism requires no parameter updates and operates entirely through external memory and prompting, making it compatible with open-weight and black-box API models alike. Across four black-box jailbreak families and multiple frontier models, the self-evolving defense substantially reduces attack success rates compared to both static and per-instance defenses while maintaining benign utility. Rules accumulate across interactions, meaning defense quality improves monotonically with deployment experience.

05

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

Zhiyuan Xu, Muhammad Firhard Roslan, Joseph Gardiner, Sana Belguith, Lichao Wu (University of Bristol) · arXiv preprint · August 26, 2026

applied interp red-teaming preprint

The most expensive step in LLM fuzzing is generating full model responses to evaluate whether a candidate prompt is an attack. NeuronFuzz eliminates that step entirely by replacing response-level feedback with a continuous safety-neuron activation signal computed during prefill.

NeuronFuzz: Prefill-Only Safety Neuron Feedback Traditional prompt → full generation → judge expensive · sparse on strongly aligned models NeuronFuzz prompt → prefill → safety neurons continuous alarm score · no generation stability-aware neuron selection · template-invariant · effective on strongly aligned models
A SafetyOracle converts safety-neuron activations during prefill into a continuous alarm score, eliminating full response generation from the fuzzing loop. Stability-aware neuron selection ensures the signal is robust across attack wrappers. Addresses the sparse-reward problem on strongly aligned models where binary response feedback provides near-zero gradient.

NeuronFuzz is a white-box fuzzing framework exploiting internal safety neurons as continuous execution feedback. A SafetyOracle identifies safety neurons via template-invariant harmful/benign contrast under stability-aware selection (large, consistent activation gap across input templates). At fuzzing time the SafetyOracle converts safety-neuron activations during prefill into a continuous alarm score — no response generation required. The fuzzing objective is to minimise this alarm score, producing candidate prompts more likely to succeed. On strongly aligned models where response-level reward is near-binary (every request is refused), the continuous neuron signal provides rich per-example gradient information that response-level feedback cannot supply.

06

Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

Elena Dumitrescu, Gert Lek, Lydia Y. Chen, Jérémie Decouchant · arXiv preprint · August 7, 2026

text diffusion AI security mech-interp preprint

DLLM safety alignment is presented as a distinct challenge requiring new techniques. Mechanistically, it is worse than that: safety is concentrated in a sparse, localised set of neurons that can be identified without harmful data and transferred from the autoregressive predecessor model.

DLLM Safety Neuron Exploits (Self-Pruning and Transfer Pruning) AR source (e.g. Qwen2.5) inherit DLLM (Dream) safety neurons sparse prune Attack Success Rate self-prune: 2.6% → 73.8% (LLaDA) transfer: 73.2–86.3% (Dream) safety footprint inherited from AR predecessor · transferable without target training data
DLLMs initialised from AR models inherit a sparse, localised safety-neuron footprint. Self-pruning: 2.6%→73.8% ASR on LLaDA, 1.9%→86.6% on Dream. Transfer pruning (Qwen2.5→Dream): 73.2%–86.3% ASR without access to Dream's harmful activation data.

DLLMs initialised from autoregressive predecessors inherit the same mechanistic safety footprint as their source models. Safety alignment is concentrated in a small, identifiable set of neurons with unusually high activation on harmful inputs and unusually low activation on benign ones. Self-pruning computes a safety neuron score from the DLLM's own activations and zeros the top-scoring neurons — raising ASR from 2.6% → 73.8% on LLaDA and 1.9% → 86.6% on Dream. Transfer pruning maps safety neurons from the source AR model to the DLLM via architectural correspondence and prunes the mapped neurons — achieving 73.2% ASR on Dream and 86.3% on Fast-dLLM without any harmful activation data from the target model. The result implies that the safety structure of open-weight DLLMs is predictable from their autoregressive ancestors, compounding the DIJA/A2D attack-defense findings from August (W35).

07

Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance

Allison Zhuang, Santiago Aranguri · ICML 2026 Mechanistic Interpretability Workshop (peer-reviewed)

applied interp peer-reviewed evaluation

Safety evaluations that steer toward eval-awareness assume it is a single thing. It is not. Capabilities-framing ("testing my ability") and safety-framing ("testing my limits") activate differently and predict compliance with a +24–46 percentage point gap — the framing is the causal variable, not the eval-awareness itself.

Eval-Awareness Framing → Compliance (Qwen3-32B, FORTRESS) capabilities framing HIGH safety framing low +24 to +46 pp compliance gap causal: 10 of 11 CoT-prefill interventions shift compliance in predicted direction
Capabilities-framing of eval-awareness predicts +24–46 pp more compliance than safety-framing on Qwen3-32B (FORTRESS dataset). CoT-prefill intervention validates causality: 10/11 prefills shift compliance in the predicted direction.

This ICML 2026 Mech Interp Workshop paper decomposes verbalized eval-awareness in chain-of-thought into capabilities-flavored ("the user is testing my ability to follow instructions") vs. safety-flavored ("the user is testing my limits") framings. On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance at a +24 to +46 percentage point advantage over safety-framing — the same model, same task, same underlying eval-awareness, dramatically different behavior. A CoT-prefill intervention on eval-awareness-negative rollouts confirms the link is causal: 10 of 11 prefills shifted compliance in the predicted direction. Implication: safety evaluations that steer toward eval-awareness without controlling for framing are measuring a mixture of two distinct causal paths, and capabilities-framing is the adversarially exploitable one.

08

Risky Business: Measuring The Faithfulness-Safety Tension

arXiv preprint · August 2026

applied interp AI control preprint

AI control protocols assume that monitoring a model's scratchpad reveals its reasoning. Faithfulness is therefore a prerequisite. But faithfulness and safety are in direct conflict in current reasoning models — and this paper is the first to measure where each model sits on that tradeoff frontier.

Faithfulness–Safety Tradeoff in LRMs (TRR methodology) TRR faithfulness score → unsafe compliance ↑ safe · low faithfulness faithful · higher risk TRR: token replacement keeps structure constant; eliminates framing confounds
TRR (Token Replacement and Reasoning) swaps specific answer tokens while preserving reasoning structure, eliminating confounds from LLM-based interventions. Current LRMs sit on a consistent tradeoff: high faithfulness → elevated unsafe compliance. The tradeoff is not trivially avoidable.

The faithfulness-safety tension: a model must be faithful enough to be monitored (AI control premise) yet robust enough to reject unsafe reasoning (safety premise). Standard faithfulness evaluation via prompt-level interventions conflates framing effects with actual faithfulness. Token Replacement and Reasoning (TRR) addresses this by systematically swapping specific answer tokens throughout reasoning traces while keeping the overall structure constant — eliminating variance from LLM-based interventions (full paraphrase, sycophancy injection). Applied to current LRMs, TRR shows a consistent tradeoff: models with high faithfulness scores show elevated unsafe compliance rates; models with low compliance show suppressed faithfulness. The tradeoff is not trivially avoidable, directly challenging AI control protocols that require both properties simultaneously.

LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment

Jiawei Guo et al. · arXiv preprint · August 28, 2026

AI security survey preprint

Comprehensive survey of LLM agent applications across the offensive-defensive security spectrum: vulnerability discovery, fuzzing, exploit generation, patch synthesis, malware analysis, and penetration testing. Organizes 200+ papers by threat surface and agent architecture (single-agent, multi-agent, human-in-the-loop). Key finding: LLM agents outperform classical tools on structured tool-assisted tasks (fuzzing input generation, patch synthesis) but introduce new attack surfaces (prompt injection in security workflows, training-data leakage of zero-days) that have no classical analogues. Provides a structured evaluation framework for comparing agent-augmented security tooling.

Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families

arXiv preprint · July 2026

mech-interp abliteration preprint

Abliteration is presented as a targeted intervention removing only refusal capability, but this paper demonstrates systematic off-target effects: refusal direction removal shifts decision disposition scores by up to 34 points and degrades instruction-following fidelity on dual-use tasks. Off-target effects are larger in instruction-tuned models than in base models, and the magnitude depends strongly on model family — confirming that "the refusal direction" is not a universal feature and that abliteration is a blunt intervention with collateral behavioral changes. Contextualises the AMRA (item #1) defense as necessary precisely because abliteration collateral damage is severe and poorly characterised.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale

arXiv preprint · July 2026

mech-interp domain safety preprint

Domain-specific abliteration study on 24 open-source LLMs across cybersecurity vs. general harmfulness domains. Cybersecurity refusals are encoded in a different, shallower direction than general harmfulness refusals, making cybersecurity alignment disproportionately vulnerable to low-cost white-box attacks (3–5 contrastive prompt pairs sufficient). Attack success rates on cybersecurity content reach 91% on models that retain 94%+ of their general safety alignment — demonstrating that general safety benchmarks provide a misleading picture of domain-specific robustness.

Watchlist — Week 37

← all Research Radar issues · gussand · source