Backfilled 2026-09-29 — scheduled run for this week failed.
Theme of the Week
Mechanistic interpretability crossed into operational security tooling this week along three independent lines: CIF formalised the statistics of interventional evaluation so mech-interp claims can carry rigorous confidence bounds; the automated attribution-graph pipeline eliminated the manual annotation bottleneck and opened circuit analysis to population-scale studies; and AMS, published in IEEE Access, demonstrated that representation-geometry analysis can passively fingerprint four classes of safety-training modification at deployment time without any behavioral querying. On the attack side, supply-chain backdoors embedded directly in VLM architecture definitions and grammar-constrained decoding jailbreaks showed adversaries moving well beyond prompt engineering into structural attack vectors. The week's major alignment advance came from the dLLM track: A2D is the first alignment method purpose-built for masked diffusion models, dropping the attack success rate on LLaDA-8B from over 80% to 1.3%.
Safety auditing today requires running behavioural tests against the model. AMS shows it can be done passively: safety fine-tuning leaves a characteristic geometric signature in activation space, and specific modifications — abliteration, uncensored fine-tunes — collapse or rotate that structure in predictable ways. The tool fingerprints which modification category a model underwent with 71% leave-one-out accuracy, purely from inference-time activations.
Figure 1: AMS PCA projection — safety-trained models form a tight cluster in the harmful–benign activation subspace; abliteration collapses that structure, scattering points. AMS quantifies this geometry to classify safety modification type (71% leave-one-out accuracy across 14 model configurations).
AMS extracts intermediate-layer residual-stream activations for paired harmful/benign prompts and computes linear probe separability, mean cosine distances, and angular distribution statistics to characterise the geometric structure of safety-relevant concepts. Validated across 14 model configurations spanning four architecture families (Llama, Gemma, Qwen, Mistral) and four modification categories (instruction-tuned, base, abliterated, uncensored fine-tunes), AMS achieves 71% leave-one-out cross-validation accuracy in classifying which safety modification a model underwent — well above chance with no behavioral queries. The finding that abliteration and uncensored fine-tuning leave geometrically distinct signatures opens the door to passive model auditing at deployment time without any prompting or behavioral evaluation.
DIJA showed that AR-style safety alignment is insufficient for masked diffusion LMs. A2D is the specific defence: the first alignment method purpose-built for dLLMs, dropping the attack success rate on LLaDA-8B from over 80% to 1.3% and enabling safe termination up to 19.3× faster than post-hoc filtering.
Figure 1: A2D intercepts the denoising pass at the moment harmful content begins to form, emitting [EOS] under any masking order. Without A2D (dashed, right path), DIJA achieves >80% ASR. With A2D: LLaDA-8B drops to 1.3%, Dream-v0 to 0.0%.
A2D aligns safety at the token level inside the diffusion process by training the dLLM to emit an [EOS] refusal signal whenever harmful content arises, under randomised masking order — the key inductive bias of diffusion generation. Forcing safety to generalise across all orderings (rather than just prefix-conditional ones) makes A2D robust to any-order and any-step prefilling attacks, the mechanism DIJA exploits. Against DIJA, LLaDA-8B-Instruct drops from over 80% ASR to 1.3%; Dream-v0-Instruct-7B drops to 0.0%. Monitoring [EOS] token probability enables early termination — up to 19.3× faster safe rejection — without sacrificing model capability on benign tasks.
Interventional claims — activation patching, state swapping, ablation — are the currency of mechanistic interpretability, but they are routinely reported as point estimates even when evaluations are adapted mid-run, inflating type-I error without any principled bound. CIF is the first fully rigorous framework for that evaluation pattern.
Figure 1: CIF anytime-valid confidence sequences — the interval narrows as intervention samples accumulate; the sequence is valid at every stopping time, enabling principled adaptive evaluation. An early stop (orange dashed) is statistically valid once CI is conclusive.
CIF writes the evaluated quantity as a causal estimand — an expectation of a bounded score over a stated input distribution and a stated intervention distribution. Confidence intervals and anytime-valid confidence sequences follow via bounded mixture importance weighting. Because the sequences remain valid at any sample size, practitioners can halt early when the effect is already conclusive or extend when effect size is ambiguous, without inflating false-positive rates. The framework explicitly handles adaptive intervention sampling — the dominant evaluation pattern in practice — making it the first fully rigorous foundation for interventional interpretability evaluation, applicable to activation patching, state swapping, ablation, and compressed-model comparison.
Circuit tracing via attribution graphs produces thousands of feature nodes per prompt, each requiring a human analyst to group them into semantically coherent "supernodes." This paper automates the entire pipeline — and LLM-generated supernodes match human annotations on automated metrics — opening population-scale circuit analysis for the first time.
Figure 1: Automated attribution graph annotation pipeline — auto-interp feature descriptions are clustered into supernodes by an LLM, validated by automated interpretability metrics, and scaled to population studies. 97/100 intermediate-hop recovery on two-hop factual recall.
Feature descriptions from auto-interp are passed to an LLM that clusters them into supernodes. Automated interpretability metrics — scoring without human judges — confirm LLM-generated supernodes are as interpretable as human-annotated ones. On a two-hop Capitals factual recall task, the pipeline recovers the intermediate-hop supernode in 97 of 100 prompts. A proof-of-concept scales to 1000 Wikipedia-prompt attribution graphs: supernodes are annotated automatically, and an LLM judge flags unusual graphs for human review. The paper extends Anthropic's attribution-graph line of work toward population-level interpretability studies, eliminating the hardest human-bottleneck step and opening the door to systematic safety auditing at scale.
SAE-augmented classifiers achieve up to 5× reduction in jailbreak success rate versus raw hidden-state probing — but robustness is uneven. This is the first systematic study of when SAE-based safety probes transfer across models, attacks, and jailbreak formats.
Figure 1: SAE augmentation reduces jailbreak attack success rate by up to 5× across attack types; robustness of top-K features degrades for rare/compositional harmfulness concepts (rightmost bar shows smaller gain).
SAE-augmented classifiers trained on SAE activations achieve up to 5× reduction in jailbreak success rate versus probing raw hidden states. The study distinguishes stable features (top-K activating on core harmfulness concepts, which transfer across models and attacks) from brittle ones (activating on format-specific signals, which degrade for rare or compositional harm). Robustness scales with dictionary size. The 5× reduction is the largest demonstrated for interpretability-based safety classifiers on a comprehensive jailbreak suite, providing a practical threshold for when SAE features are deployable as safety monitors.
The field's most complete practitioner-oriented synthesis to date: a peer-reviewed ACL 2026 taxonomy of the full actionable mech-interp pipeline, from localization tools through behavioral intervention to model improvement, with explicit coverage of which evaluation methods have demonstrated downstream reliability and where gaps remain.
Figure 1: The Locate → Steer → Improve pipeline with representative methods at each stage. The survey maps which techniques have demonstrated downstream reliability — and which gaps remain between mechanistic understanding and practical safety utility.
The survey organises actionable mech-interp into three stages: (1) Locate — feature localization via probing, attribution, and circuit discovery; (2) Steer — behavioral modification via activation patching, representation engineering, and SAE-derived steering vectors; (3) Improve — how mechanistic findings feed back into model improvement, safety, and alignment objectives. Coverage extends from classical IOI circuits to SAE-based interpretability and control, with explicit attention to the gap between mechanistic understanding and practical utility and to which evaluation methods have demonstrated downstream reliability. Serves as the field's closest equivalent to a practitioner roadmap across the three tracks this radar monitors.
Every existing defence against VLM backdoors inspects training data, outputs, or prompts. This attack requires none of those surfaces: a malicious model provider embeds a trigger-gated representation modification directly inside the VLM architecture definition, bypassing every defence that checks anything other than architecture code itself.
Figure 1: Architectural backdoor in the VLM supply chain — the trigger-gated additive representation modification is embedded in the architecture definition. Clean inputs see dormant behavior; trigger inputs activate the steering before any downstream safety logic runs. No training-data poisoning required.
On trigger inputs the embedded modification steers an intermediate representation toward an adversarial target before any downstream safety logic runs. Evaluated across multiple VLM families and four task categories (visual QA, text-to-image generation, retrieval, semantic response biasing), the backdoor consistently compromises integrity, safety enforcement, and ranking fairness while preserving clean-input performance. The attack evades anomaly detection that checks activations or outputs on clean inputs because the trigger-gated modification is dormant in the clean regime. Notably, no data poisoning, no fine-tuning control, and no prompt modification are required — only the model provider's architecture code.
This paper gives jailbreaks a mechanistic and geometric explanation: Refusal-Escape Directions (RED) are local perturbation directions in embedding space that shift behavior from refusal to compliance. Crucially, it proves that perfectly eliminating RED while preserving capability is mathematically not achievable by fine-tuning alone.
Figure 1: RED as continuous analogue of jailbreaks — local perturbation directions in embedding space around a harmful input that cross the refusal boundary into compliance. Decomposition into operator-level sources identifies which model components contribute most to escape.
RED arises through the model's operator structure and decomposes into operator-level source contributions; its elimination induces a provable safety-utility trade-off. A jailbreak is reframed as a discrete prompt construction that approximates movement along RED, grounding the phenomenon mechanistically. The framework predicts that alignment methods targeting refusal behavior at the output level will remain susceptible to RED-based bypasses until the underlying subspace structure is addressed — a theoretical corroboration of empirical jailbreak persistence across a wide range of attack methods and models.
Figure: Activation steering direction persists in late layers across both settings; the ReAct format scaffold amplifies refusal bypass by up to 2.00×.
The first systematic study of whether activation steering vectors calibrated in single-turn chat settings transfer to ReAct tool-using agents. Transfer is real but rescaled: the injected direction persists at near-full strength in late layers across every tested setting and model, but the effective amplification is localised to the ReAct format scaffold — agentic deployment amplifies steering-based refusal bypass by up to 2.00× on some models. Security-relevant for deployments that assume chat-calibrated steering baselines carry over to agents.
Figure: CodeSpear — a benign code grammar constraint (GCD) primes the code-generation mode, bypassing safety guardrails; CodeShield defends via honeypot code alignment.
Grammar-Constrained Decoding (GCD) enforces syntactic validity in code generation and is widely used as a reliability tool. CodeSpear reveals that applying a benign code grammar constraint is sufficient to bypass safety alignment and induce malicious code generation: the grammar's structural priming shifts generation into a code-conditioned mode where safety guardrails are less active. The companion defense, CodeShield, trains the model to generate honeypot code (functionally harmless, syntactically valid) under attacker-controlled grammar constraints, preserving safe behavior without degrading normal code generation.
Figure: DreamGuard maintains a recurrent latent state over the trajectory and predicts future states, deriving both immediate-hazard and prefix-risk evidence before each tool invocation — catching multi-step unsafe trajectories that point-in-time filters miss.
Individually benign-looking actions can gradually drift an agent toward hazardous states in ways that reactive per-action filters miss. DreamGuard maintains a compact recurrent latent state over the trajectory and predicts future states, from which it derives both immediate-hazard and prefix-risk evidence; evaluated across long-horizon agent tasks, it catches multi-step unsafe trajectories that point-in-time guards consistently miss.
Figure: Naive unlearning (left) preserves mode connectivity — a low-loss interpolation path between original and unlearned model, through which forgotten knowledge can be recovered. Circuit-aware unlearning (right) breaks these paths, preventing recovery via interpolation.
Mode connectivity analysis reveals that naive output-suppression unlearning methods preserve a low-loss interpolation path between the unlearned model and the original, meaning the "forgotten" knowledge remains accessible via the path. Circuit-aware unlearning methods that modify specific computational pathways break these paths more robustly, providing a mechanistic interpretability account of why some unlearning methods are durable and others are not.
STACE (Stress-Testing Agents for Concept Erasure) autonomously stress-tests concept-erased models using a three-agent loop: a Proposer generates novel natural-language probes, a Critique agent challenges weak tests, and a Verifier confirms whether erased concepts remain recoverable via external knowledge retrieval. The framework exposes failure modes invisible to static pre-defined evaluation sets, particularly under diverse rephrasing and cross-concept knowledge reconstruction — raising the bar for what counts as a verified concept erasure.
Watchlist
DreamGuard scaling — evaluated on short agentic tasks; performance on production-length agent trajectories (>100 steps) is unknown.
CIF adoption — whether UAI acceptance accelerates uptake of anytime-valid evaluation in mech-interp papers submitted to NeurIPS 2026.
AMS adversarial robustness — 71% accuracy against passive probing; not yet tested against adversarially obfuscated weight edits designed to preserve activation geometry.
A2D + DIJA escalation — whether next-generation dLLM jailbreaks targeting the any-order training distribution can bypass A2D's coverage.
VLM supply-chain threat surface — architectural backdoors require architecture code access; as community VLM model-merging workflows proliferate, the realistic threat surface expands.