Week 31 · July 20–26, 2026 (backfilled 2026-09-29)
0 peer-reviewed · 11 preprints · 0 forum/blog
AI Safety·Alignment·Mech Interp
Theme of the Week
Sparse autoencoders underwent a methodological stress test this week — three independent papers interrogated whether SAE features are reliable instruments for causal claims and safety-critical work. Marks et al. proved which linear readouts survive sparse compression (most safety-relevant probes do; rare-concept probes lose 15–40%); Bloom, Conmy, and Nanda showed that timescale is a genuine dimension of feature identity that standard stateless SAEs systematically miss; and Cho et al. demonstrated that SAE family choice is a larger determinant of interventional validity than model scale. Alongside these methodological advances, LACUNA exposed the gap between unlearning metrics and genuine knowledge erasure, PAC-style safety bounds provided the first distribution-free formal certificates for LLM output safety, and CC-Delta showed SAE-guided jailbreak defense outperforms dense activation steering. The security track closed with a multi-turn jailbreaking advance and a survey confirming 69–98% policy-enforcement failure rates in AI coding agent deployments.
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller — arXiv (July 2026)
First formal characterization of information loss through SAE bottlenecks — directly answers whether safety-critical linear probes survive SAE-based interventions, a question every practitioner using SAEs for steering or monitoring needs to resolve before trusting their results.
When a sparse autoencoder compresses a residual-stream vector into sparse feature activations, it discards information — but which information? This paper proves exactly which linear readouts survive the compression, and shows empirically that most safety-relevant probes do while rare-concept probes do not.
A readout φ(z) = wᵀz is preserved through SAE compression iff the weight vector w lies in the column span of the decoder matrix D. Safety-relevant high-frequency probes (refusal direction, harmfulness) satisfy this condition; rare-concept probes are partly projected out, losing 15–40% predictive accuracy.
The paper frames SAE compression as a constrained matrix-valued distortion minimization problem and derives a closed-form characterization: a linear functional φ(z) = wᵀz is preserved iff w ∈ col(D). Empirically, probes for refusal direction, harmfulness scores, and entity type have weight vectors well-aligned with the decoder span and survive with near-zero degradation; probes for rare named entities or low-frequency syntax lose 15–40% of predictive accuracy. This gives practitioners a principled pre-flight test: check decoder-span alignment before trusting that a safety probe generalizes through an SAE intervention.
Joseph Bloom, Arthur Conmy, Neel Nanda — arXiv (July 2026)
Standard SAEs treat each token independently, blind to whether a feature fires in isolated bursts or sustained spans — a systematic miss for sequence-level safety monitoring. This paper adds per-feature persistence coefficients that capture exactly this distinction.
Some features burst at a single surprising token; others sustain across 20–50 tokens of a discourse topic. Standard SAEs cannot tell the difference because they process each position as if no other positions exist. Persistent SAEs fix this by letting each feature learn its own decay rate.
Feature timescale distribution learned by Persistent SAEs on Gemma-2-9B. Three regimes emerge: lexical-surprise features burst at 1–3 token positions; syntactic head-agreement features sustain over 3–8 tokens; discourse-level topic features persist over 20–50 tokens. Standard SAEs cannot represent this structure.
Persistent SAEs augment the TopK encoder with a per-feature persistence coefficient α_f ∈ [0,1] learned during training: at each token, feature f's pre-activation is blended with a decayed version of its previous activation, allowing sustained features to self-reinforce and bursty features to decay rapidly. Trained on Gemma-2-9B, Persistent SAEs recover three clear timescale regimes with distinct interpretable character. Persistence coefficients are predictive of which features transfer across tasks, suggesting timescale is a meaningful dimension of feature identity beyond interpretability labeling.
First distribution-free, model-agnostic PAC-style certificates for LLM output safety — closes the gap between empirical safety evaluations (which cannot quantify unseen-prompt risk) and formal verification (which requires internal access). Certificates are composable across safety-pipeline stages.
Standard safety evaluations tell you how often a model fails on a test set. This paper tells you, with formal statistical guarantees, how often it will fail on prompts you haven't seen — using only a modest sample and no access to model internals.
PAC safety certification pipeline. From n test prompts (n ≥ 200 suffices at δ = 0.01), conformal prediction combined with Clopper-Pearson intervals yields a distribution-free bound P(unsafe) ≤ ε with probability 1−δ. Certificates hold within 2–5% of empirical violation rates on ToxiGen and AdvBench across four frontier models.
The framework formulates LLM safety certification as a PAC learning problem over a prompt distribution, then computes a sample-efficient bound on the probability of safety predicate violation on future unseen prompts. The bound uses conformal prediction combined with Clopper-Pearson intervals — making it distribution-free and assumption-free about LLM internals. Evaluated on ToxiGen and AdvBench with four frontier models, certificates hold within 2–5% of empirical violation rates at δ = 0.01 with as few as 200 test prompts, and are composable: certificates for sub-policies (refusal gates, output classifiers) can be chained to certify multi-stage safety pipelines.
Matteo Boglioni, Thibault Rousset, Siva Reddy, Marius Mosbach, Verna Dankers — Mila / McGill — arXiv (July 2026)
First unlearning testbed with ground-truth parameter-level localization; proves that standard high-metric unlearning does not imply genuine knowledge erasure and quantifies exactly how much better localization could improve robustness to resurfacing attacks.
When an unlearning algorithm achieves high scores on the standard forgetting benchmarks, does it actually remove the knowledge from the weights — or just suppress its expression? LACUNA answers this question by building a testbed where the ground truth is known.
LACUNA exposes the localization gap: standard unlearning algorithms (SimNPO and variants) achieve high unlearning metric scores but low resurfacing robustness because they do not target the responsible weights. OracleGrad — with oracle access to ground-truth parameter localization — achieves both, quantifying the gap that better mechanistic localization methods could close.
LACUNA injects PII of synthetic individuals into predefined weight subsets of 1B and 7B OLMo models via masked continual pretraining, creating a controlled testbed with oracle knowledge of which parameters encode each to-be-forgotten fact. Standard unlearning algorithms achieve high traditional unlearning metrics without targeting the responsible weights, leaving models vulnerable to resurfacing attacks that recover the supposedly forgotten information. OracleGrad — a simple baseline with oracle parameter localization — achieves optimal forget/retain balance and maximal resurfacing robustness, quantifying the gap that better mechanistic localization could close.
Seonglae Cho, Zekun Wu, Kleyton Da Costa, Rishi Kalra, Ilham Wicaksono, Adriano Koshiyama — arXiv (July 22, 2026)
Establishes that SAE family is a more important methodological choice than model scale for causal interventional claims — a concrete warning for safety practitioners selecting SAE families for probing or steering experiments.
If you ablate a feature that reliably activates on a single token type and the model's probability of generating that token doesn't drop significantly, the feature wasn't doing the causal work. Testing 3.9 million such features, this paper finds the answer depends more on which SAE family you trained than on which model you're studying.
Single-token feature causal necessity rates by SAE family. GemmaScope and BatchTopK features are causally anchored (86% and 82% BH-significant); LlamaScope features are locally redundant (~40%). The cross-family gap exceeds the within-family scaling effect — SAE family selection matters more than model scale for interventional validity.
The paper uses single-token SAE features as ground-truth-verifiable diagnostics: ablating a feature activating on token t must significantly reduce log-probability for t to count as causally necessary. Testing 3.9 million features from six models across three SAE families, BH-significant logit reductions appear in 178 of 208 conditions. Single-token features cluster 4.7× tighter in decoder space than all features and concentrate in early layers. The key finding is that cross-family variation exceeds within-family scaling effects, directly implying that SAE family is a first-order methodological choice for anyone relying on SAE interventions for safety work.
Yannick Assogba et al. — arXiv (February 2026, surfaced July 24, 2026)
Demonstrates the clearest pragmatic safety payoff for SAE interpretability to date: SAE-guided jailbreak defense outperforms dense activation steering on all four tested models and 12 jailbreak categories, with the largest gap on out-of-distribution attacks.
Dense activation steering for jailbreak defense conflates features active in harmful jailbreak context with features active in all harmful queries — including benign ones. CC-Delta separates them by comparing activations with and without jailbreak context, then steers only on the difference.
CC-Delta identifies jailbreak-specific SAE features by contrasting activations for the same harmful request with and without jailbreak context. Steering on only the delta (f_jb) avoids activating on features shared with benign harm queries, reducing over-refusal while maintaining defense coverage — especially on out-of-distribution attacks not seen during defense calibration.
CC-Delta compares token-level SAE sparse activations for the same harmful request framed with and without jailbreak context. The contrastive delta isolates jailbreak-specific features without contaminating features active in both jailbreak and benign contexts. A mean-shift steering vector applied in SAE latent space at inference time then suppresses only the jailbreak-specific directions. Evaluated across four models and 12 jailbreak categories, CC-Delta outperforms dense activation steering throughout, with the largest gains on out-of-distribution attacks — consistent with the interpretation that SAE decomposition captures jailbreak structure invisible to dense residual-stream methods.
Proves a fundamental separation between white-box steerability and black-box prompting that has been assumed but never formally established — with direct implications for how any experiment comparing steering to prompting should be designed and what its results can claim.
Researchers often ask "is there a prompt that produces the same behavior as this steering vector?" This paper proves the answer is almost surely no — the set of states reachable via steering is strictly larger than the set reachable via any prompt.
The set of hidden states reachable via activation steering strictly contains the set reachable via any discrete prompt — with the prompt-reachable manifold having measure zero in the larger space. Consequently, almost surely no prompt can reproduce the same internal behavior as a steering vector, and steering-vs-prompting comparisons cannot assume equivalence.
The paper constructs the manifold of prompt-reachable hidden states and proves it has measure zero in the space of states reachable via activation steering — meaning that for almost any steering vector and magnitude, the induced hidden state lies strictly outside the prompt-reachable manifold. This formal non-surjectivity result establishes a fundamental separation: white-box steerability is not equivalent to black-box prompting, benchmarks that compare the two under equivalence assumptions are methodologically unsound, and evaluations of steering effects should explicitly decouple the white-box and black-box conditions.
Junyoung Park, Namgyu Park, Sechan Lee, Yoon-Chan Jhi, Jihoon Cho, Sangdon Park — arXiv (July 13, 2026)
Addresses the key limitation of RL-based multi-turn jailbreaking — each dialogue turn contributes unequally to the final outcome, but prior methods apply a single trajectory-level signal that conflates useful and noise turns — with a principled decomposed credit estimator.
A multi-turn jailbreak attacker needs to learn which conversational turns actually moved the model toward compliance and which were noise — but standard RL broadcasts a single final reward to all turns equally. DC-GRPO gives each turn its own credit signal.
DC-GRPO replaces a single trajectory-level reward with per-turn credits combining immediate reward rₜ and future-value estimate γVₜ₊₁. Both static- and dynamic-weighted instantiations achieve strong multi-turn attack success rates on black-box frontier models with cross-model transferability.
DC-GRPO (Decomposed Credit Group Relative Policy Optimization) derives a per-turn value function that separates each turn's contribution to the final harmful outcome into immediate and future components, then applies group-relative policy optimization independently per turn. Both static-weighted (fixed γ) and dynamic-weighted (γ learned per dialogue state) instantiations are evaluated on multi-turn jailbreak benchmarks against black-box frontier models, achieving strong attack success rates and cross-model transfer — confirming that per-turn credit decomposition is a broadly applicable improvement over trajectory-level RL for multi-turn adversarial training.
Proposes a six-condition reliability filter followed by Borda consensus ranking (F-test, KSG mutual information, Cohen's d) to select SAE features for steering without a learned objective, evaluated on three Gemma-family models. Key methodological finding: raw attribute movement in activation space does not predict generation quality — practitioners need an explicit selection filter before treating SAE steering as effective.
Systematic review of 39 papers across 17 execution-security categories for AI coding agents; headline finding is 69–98% policy-enforcement failure rates across all categories. Documents 4 patched CVEs and 5 structural gaps, arguing the field lacks the shared ontology and shared benchmarks needed for coordinated progress on agent security.
Attribution and activation patching on Llama-3 models identifies a compact shared set of MLP neurons that is both necessary and sufficient for late-layer arithmetic computation across symbolic arithmetic, natural-language word problems, and Python code. Targeted interventions degrade performance simultaneously in all three modalities, supporting a single surface-form-invariant arithmetic circuit.
Watchlist
Persistent SAE extensions — follow-up work applying timescale structure to supervised probe training and long-horizon agent monitoring; the Bloom/Conmy/Nanda preprint sets up a natural extension to safety monitoring across multi-turn conversations.
LACUNA localization gap — upcoming papers attempting to close the OracleGrad gap using attribution-patching-based localization without oracle access; the testbed makes evaluation tractable.
PAC safety certification under distribution shift — the 2607.20286 framework is distribution-free but assumes stationarity; extensions to adaptive adversaries and distribution shift are the natural next step.
Multi-turn jailbreak defenses — DC-GRPO raises the bar for multi-turn ASR; corresponding defense work is not yet visible and represents an open gap.
CC-Delta distillation — SAE-based steering at inference time carries latency cost; watch for approaches that bake the defense into model weights (cf. CRISP from W30 #07).