The week's dominant signal is the geometry of the refusal mechanism under adversarial pressure. Three independent papers — Abliteration Mitigation via Refusal Aliases, Decided Upstream Written Late, and Tripwire — simultaneously attack the same question from defense, circuit-tracing, and statistical-certification angles, revealing that the refusal signal is both architecturally fragmented (harm is detected mid-network in a language-invariant direction, but refusal is written late by a localised circuit) and geometrically obfuscatable (the refusal direction is easily extracted and projected out unless actively aliased). On the security side, two Aug 26 submissions bring automated red-teaming and adaptive defense into production range: NeuronFuzz eliminates response-generation from the fuzzing loop by using safety-neuron activations as continuous prefill-only feedback, while a self-evolving defense abstracts attack patterns into method-level rules that generalise across attack families. Text diffusion LMs remain a live mechanistic attack surface: safety neurons in DLLMs are sparse, localised, and transferable from autoregressive predecessors, with self-pruning raising ASR from 2.6% → 73.8% on LLaDA. An ICML 2026 Mech Interp Workshop result caps the week by decomposing eval-awareness into capabilities-framing vs. safety-framing, showing the former predicts compliance with a +24–46 pp gap — with direct implications for red-teaming and monitoring pipeline design.
Abliteration — removing the refusal direction via projection — has become a commodity attack requiring only a handful of contrastive prompts. AMRA is the first defense that addresses the root cause: the extractability of the refusal direction itself.
AMRA replaces refusal-inducing activations with random aliases in writer matrices; reader matrices corrected to preserve behavior. Contrastive extraction finds the alias, not the true direction. Llama-3-8B: +2.16 post-abliteration refusal score, <0.5 pp MMLU cost.
Abliteration works by extracting the refusal direction r̂ from contrastive harmful/benign prompt pairs, then projecting weight matrices orthogonal to r̂. AMRA applies rank-k updates to residual-stream writer matrices to replace the true refusal-inducing activation with a random alias, then corrects downstream reader matrices to preserve the original model behavior. The net effect: the contrastive extraction procedure finds the alias (which no longer encodes refusal) rather than the true direction, so projection-based removal has no effect on actual refusal behavior. On Llama-3-8B, AMRA improves post-abliteration refusal scores by 2.16 points over the undefended baseline with less than 0.5 percentage points of MMLU degradation.
Safety gaps in multilingual models look like failures to detect harm. They are not. Harm is detected in a language-invariant direction mid-network — but refusal is written by a separate, late, localised circuit whose architecture makes cross-lingual transfer expensive in one direction and cheap in the other.
Harm detection is language-invariant at L11 (cosine ≈ 0.9, English vs. Indic). The refusal write is late, localised to a MoE writer + attention-opposer circuit, and geometrically orthogonal to detection. Damping the opposer is cheap; amplifying the writer hits a cost wall with context length.
Evaluated on Sarvam-30B, an Indic-multilingual mixture-of-experts model. Harm is encoded as a direction nearly language-invariant in mid-network (English–Indic cosine ≈ 0.9 at layer 11); steering this direction causally controls refusal output. However, this detection direction is geometrically orthogonal to the direction that actually writes the refusal token — the write is late in the network and assembled across generation. Circuit attribution localises the write to a specific MoE writer head held in check by an attention opposer. Key interventional asymmetry: damping the opposer is cheap and reliably produces refusal; amplifying the writer hits a cost wall (required intervention strength grows with context length). Cross-lingual safety gaps persist not because harm isn't detected, but because the write signal doesn't transfer across languages — explaining why naive fine-tuning on English refusal examples fails to close the gap.
Every existing jailbreak defense either perturbs all requests (utility cost) or requires a separate training run (deployment cost). Tripwire identifies a certified compact set of safety neurons via hypothesis testing and patches only those neurons — only when needed.
Per-neuron hypothesis tests under Benjamini-Hochberg FDR control identify a compact safety-neuron set S. At inference time the patch fires only on alarmed inputs — zero cost on benign. Tested on 4 models × 4 attack types (GCG, AmpleGCG, AutoDAN, Jailbreak-R1).
Tripwire identifies safety neurons using per-neuron hypothesis tests comparing activations on harmful vs. benign inputs under Benjamini-Hochberg FDR control, with utility-specific criteria excluding neurons that also fire on benign requests. At inference time, when a safety alarm fires (computed from safety-neuron activations during prefill), Tripwire patches only the identified neurons to reinstate their refusal-inducing activation pattern — all other neurons are unmodified, and benign inputs incur zero intervention cost. Evaluated on Llama-2-7B, Llama-3.1-8B, Qwen2.5-7B, and Qwen2.5-32B against GCG, AmpleGCG, AutoDAN, and Jailbreak-R1, compared against RepE, TraceRouter, LED, and DELMAN baselines. Tripwire matches or exceeds state-of-the-art safety defense with lower MT-Bench degradation than weight-editing methods.
Tongyan Hu, Bryan Hooi (National University of Singapore) · arXiv preprint · August 26, 2026
AI securityjailbreak defensepreprint
Static defenses fail under adaptive attackers because new attack wrappers require manual updates. This framework instead learns from each failure: it abstracts the structural pattern of a successful attack into a method-level rule that generalises across the entire attack family — no parameter updates required.
On attack success, the framework abstracts the structural wrapper into a method-level rule stored in persistent cross-interaction memory. Retrieval via semantic similarity at inference time generalises across the entire attack family. Operates through external memory and prompting only — no parameter updates; compatible with black-box APIs.
Because rules are method-level (capturing the structural wrapper — role-play, code-transformation, multi-step indirection — rather than the harmful topic), one induced rule generalises across an entire attack family. The mechanism requires no parameter updates and operates entirely through external memory and prompting, making it compatible with open-weight and black-box API models alike. Across four black-box jailbreak families and multiple frontier models, the self-evolving defense substantially reduces attack success rates compared to both static and per-instance defenses while maintaining benign utility. Rules accumulate across interactions, meaning defense quality improves monotonically with deployment experience.
Zhiyuan Xu, Muhammad Firhard Roslan, Joseph Gardiner, Sana Belguith, Lichao Wu (University of Bristol) · arXiv preprint · August 26, 2026
applied interpred-teamingpreprint
The most expensive step in LLM fuzzing is generating full model responses to evaluate whether a candidate prompt is an attack. NeuronFuzz eliminates that step entirely by replacing response-level feedback with a continuous safety-neuron activation signal computed during prefill.
A SafetyOracle converts safety-neuron activations during prefill into a continuous alarm score, eliminating full response generation from the fuzzing loop. Stability-aware neuron selection ensures the signal is robust across attack wrappers. Addresses the sparse-reward problem on strongly aligned models where binary response feedback provides near-zero gradient.
NeuronFuzz is a white-box fuzzing framework exploiting internal safety neurons as continuous execution feedback. A SafetyOracle identifies safety neurons via template-invariant harmful/benign contrast under stability-aware selection (large, consistent activation gap across input templates). At fuzzing time the SafetyOracle converts safety-neuron activations during prefill into a continuous alarm score — no response generation required. The fuzzing objective is to minimise this alarm score, producing candidate prompts more likely to succeed. On strongly aligned models where response-level reward is near-binary (every request is refused), the continuous neuron signal provides rich per-example gradient information that response-level feedback cannot supply.
Elena Dumitrescu, Gert Lek, Lydia Y. Chen, Jérémie Decouchant · arXiv preprint · August 7, 2026
text diffusionAI securitymech-interppreprint
DLLM safety alignment is presented as a distinct challenge requiring new techniques. Mechanistically, it is worse than that: safety is concentrated in a sparse, localised set of neurons that can be identified without harmful data and transferred from the autoregressive predecessor model.
DLLMs initialised from AR models inherit a sparse, localised safety-neuron footprint. Self-pruning: 2.6%→73.8% ASR on LLaDA, 1.9%→86.6% on Dream. Transfer pruning (Qwen2.5→Dream): 73.2%–86.3% ASR without access to Dream's harmful activation data.
DLLMs initialised from autoregressive predecessors inherit the same mechanistic safety footprint as their source models. Safety alignment is concentrated in a small, identifiable set of neurons with unusually high activation on harmful inputs and unusually low activation on benign ones. Self-pruning computes a safety neuron score from the DLLM's own activations and zeros the top-scoring neurons — raising ASR from 2.6% → 73.8% on LLaDA and 1.9% → 86.6% on Dream. Transfer pruning maps safety neurons from the source AR model to the DLLM via architectural correspondence and prunes the mapped neurons — achieving 73.2% ASR on Dream and 86.3% on Fast-dLLM without any harmful activation data from the target model. The result implies that the safety structure of open-weight DLLMs is predictable from their autoregressive ancestors, compounding the DIJA/A2D attack-defense findings from August (W35).
Safety evaluations that steer toward eval-awareness assume it is a single thing. It is not. Capabilities-framing ("testing my ability") and safety-framing ("testing my limits") activate differently and predict compliance with a +24–46 percentage point gap — the framing is the causal variable, not the eval-awareness itself.
Capabilities-framing of eval-awareness predicts +24–46 pp more compliance than safety-framing on Qwen3-32B (FORTRESS dataset). CoT-prefill intervention validates causality: 10/11 prefills shift compliance in the predicted direction.
This ICML 2026 Mech Interp Workshop paper decomposes verbalized eval-awareness in chain-of-thought into capabilities-flavored ("the user is testing my ability to follow instructions") vs. safety-flavored ("the user is testing my limits") framings. On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance at a +24 to +46 percentage point advantage over safety-framing — the same model, same task, same underlying eval-awareness, dramatically different behavior. A CoT-prefill intervention on eval-awareness-negative rollouts confirms the link is causal: 10 of 11 prefills shifted compliance in the predicted direction. Implication: safety evaluations that steer toward eval-awareness without controlling for framing are measuring a mixture of two distinct causal paths, and capabilities-framing is the adversarially exploitable one.
AI control protocols assume that monitoring a model's scratchpad reveals its reasoning. Faithfulness is therefore a prerequisite. But faithfulness and safety are in direct conflict in current reasoning models — and this paper is the first to measure where each model sits on that tradeoff frontier.
TRR (Token Replacement and Reasoning) swaps specific answer tokens while preserving reasoning structure, eliminating confounds from LLM-based interventions. Current LRMs sit on a consistent tradeoff: high faithfulness → elevated unsafe compliance. The tradeoff is not trivially avoidable.
The faithfulness-safety tension: a model must be faithful enough to be monitored (AI control premise) yet robust enough to reject unsafe reasoning (safety premise). Standard faithfulness evaluation via prompt-level interventions conflates framing effects with actual faithfulness. Token Replacement and Reasoning (TRR) addresses this by systematically swapping specific answer tokens throughout reasoning traces while keeping the overall structure constant — eliminating variance from LLM-based interventions (full paraphrase, sycophancy injection). Applied to current LRMs, TRR shows a consistent tradeoff: models with high faithfulness scores show elevated unsafe compliance rates; models with low compliance show suppressed faithfulness. The tradeoff is not trivially avoidable, directly challenging AI control protocols that require both properties simultaneously.
Jiawei Guo et al. · arXiv preprint · August 28, 2026
AI securitysurveypreprint
Comprehensive survey of LLM agent applications across the offensive-defensive security spectrum: vulnerability discovery, fuzzing, exploit generation, patch synthesis, malware analysis, and penetration testing. Organizes 200+ papers by threat surface and agent architecture (single-agent, multi-agent, human-in-the-loop). Key finding: LLM agents outperform classical tools on structured tool-assisted tasks (fuzzing input generation, patch synthesis) but introduce new attack surfaces (prompt injection in security workflows, training-data leakage of zero-days) that have no classical analogues. Provides a structured evaluation framework for comparing agent-augmented security tooling.
Abliteration is presented as a targeted intervention removing only refusal capability, but this paper demonstrates systematic off-target effects: refusal direction removal shifts decision disposition scores by up to 34 points and degrades instruction-following fidelity on dual-use tasks. Off-target effects are larger in instruction-tuned models than in base models, and the magnitude depends strongly on model family — confirming that "the refusal direction" is not a universal feature and that abliteration is a blunt intervention with collateral behavioral changes. Contextualises the AMRA (item #1) defense as necessary precisely because abliteration collateral damage is severe and poorly characterised.
Domain-specific abliteration study on 24 open-source LLMs across cybersecurity vs. general harmfulness domains. Cybersecurity refusals are encoded in a different, shallower direction than general harmfulness refusals, making cybersecurity alignment disproportionately vulnerable to low-cost white-box attacks (3–5 contrastive prompt pairs sufficient). Attack success rates on cybersecurity content reach 91% on models that retain 94%+ of their general safety alignment — demonstrating that general safety benchmarks provide a misleading picture of domain-specific robustness.
Watchlist — Week 37
NeurIPS 2026 Interpretability as a Science Workshop — notifications due September 29; preprints will surface shortly after. Watch for extensions of J-Lens/Koopman identifiability to larger models (W35 item #2).
Abliteration countermeasures — four papers this week (2608.18093, 2608.14392, 2607.17427, 2607.02714) signal an active attack/defense literature. Expect systematic comparison and an adversarial red-team of AMRA.
Decided Upstream replications — the detection-writing dissociation (2608.08032) was shown in one MoE model. Cross-architecture replications (encoder-only MoE, models without the opposer circuit) are the critical next step.
EMNLP 2026 camera-ready — "Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning" confirmed accepted; preprint expected this week.
Goodfire BSF → language models — block-sparse featurizers (2606.25234, W35) demonstrated on vision; language-model extension anticipated.
SHADOWMASK (carried from W30/W35) — adversarial mask scheduling for dLLMs; still not on arXiv.