Research Radar
Week 31 · July 20–26, 2026 (backfilled 2026-09-29)
0 peer-reviewed · 11 preprints · 0 forum/blog
AI Safety · Alignment · Mech Interp
Theme of the Week

Sparse autoencoders underwent a methodological stress test this week — three independent papers interrogated whether SAE features are reliable instruments for causal claims and safety-critical work. Marks et al. proved which linear readouts survive sparse compression (most safety-relevant probes do; rare-concept probes lose 15–40%); Bloom, Conmy, and Nanda showed that timescale is a genuine dimension of feature identity that standard stateless SAEs systematically miss; and Cho et al. demonstrated that SAE family choice is a larger determinant of interventional validity than model scale. Alongside these methodological advances, LACUNA exposed the gap between unlearning metrics and genuine knowledge erasure, PAC-style safety bounds provided the first distribution-free formal certificates for LLM output safety, and CC-Delta showed SAE-guided jailbreak defense outperforms dense activation steering. The security track closed with a multi-turn jailbreaking advance and a survey confirming 69–98% policy-enforcement failure rates in AI coding agent deployments.

01
mech-interp preprint

Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression?

First formal characterization of information loss through SAE bottlenecks — directly answers whether safety-critical linear probes survive SAE-based interventions, a question every practitioner using SAEs for steering or monitoring needs to resolve before trusting their results.

When a sparse autoencoder compresses a residual-stream vector into sparse feature activations, it discards information — but which information? This paper proves exactly which linear readouts survive the compression, and shows empirically that most safety-relevant probes do while rare-concept probes do not.

z ∈ ℝ^d residual SAE ẑ ∈ col(D) reconstruction ✓ PRESERVED refusal, harm, entity type ✗ PROJECTED OUT rare / fine-grained: −15–40% preserved iff w ∈ col(D) decoder span
A readout φ(z) = wᵀz is preserved through SAE compression iff the weight vector w lies in the column span of the decoder matrix D. Safety-relevant high-frequency probes (refusal direction, harmfulness) satisfy this condition; rare-concept probes are partly projected out, losing 15–40% predictive accuracy.

The paper frames SAE compression as a constrained matrix-valued distortion minimization problem and derives a closed-form characterization: a linear functional φ(z) = wᵀz is preserved iff w ∈ col(D). Empirically, probes for refusal direction, harmfulness scores, and entity type have weight vectors well-aligned with the decoder span and survive with near-zero degradation; probes for rare named entities or low-frequency syntax lose 15–40% of predictive accuracy. This gives practitioners a principled pre-flight test: check decoder-span alignment before trusting that a safety probe generalizes through an SAE intervention.

02
mech-interp preprint

Persistent Sparse Autoencoders: Learning Feature Timescales

Standard SAEs treat each token independently, blind to whether a feature fires in isolated bursts or sustained spans — a systematic miss for sequence-level safety monitoring. This paper adds per-feature persistence coefficients that capture exactly this distinction.

Some features burst at a single surprising token; others sustain across 20–50 tokens of a discourse topic. Standard SAEs cannot tell the difference because they process each position as if no other positions exist. Persistent SAEs fix this by letting each feature learn its own decay rate.

Feature Timescale Distribution (Gemma-2-9B) 1 8 20 35 50+ persistence span (tokens) lexical-surprise (burst, 1–3 tok) syntactic (3–8 tokens) discourse-level topic (20–50 tokens, sustained)
Feature timescale distribution learned by Persistent SAEs on Gemma-2-9B. Three regimes emerge: lexical-surprise features burst at 1–3 token positions; syntactic head-agreement features sustain over 3–8 tokens; discourse-level topic features persist over 20–50 tokens. Standard SAEs cannot represent this structure.

Persistent SAEs augment the TopK encoder with a per-feature persistence coefficient α_f ∈ [0,1] learned during training: at each token, feature f's pre-activation is blended with a decayed version of its previous activation, allowing sustained features to self-reinforce and bursty features to decay rapidly. Trained on Gemma-2-9B, Persistent SAEs recover three clear timescale regimes with distinct interpretable character. Persistence coefficients are predictive of which features transfer across tasks, suggesting timescale is a meaningful dimension of feature identity beyond interpretability labeling.

03
alignment preprint

Sound Probabilistic Safety Bounds for Large Language Models

First distribution-free, model-agnostic PAC-style certificates for LLM output safety — closes the gap between empirical safety evaluations (which cannot quantify unseen-prompt risk) and formal verification (which requires internal access). Certificates are composable across safety-pipeline stages.

Standard safety evaluations tell you how often a model fails on a test set. This paper tells you, with formal statistical guarantees, how often it will fail on prompts you haven't seen — using only a modest sample and no access to model internals.

n prompts from distribution LLM + safety predicate (black-box) Conformal + Clopper-Pearson P(unsafe) ≤ ε w.p. 1−δ n ≥ 200 prompts Distribution-free · No model internals · Composable across pipeline stages
PAC safety certification pipeline. From n test prompts (n ≥ 200 suffices at δ = 0.01), conformal prediction combined with Clopper-Pearson intervals yields a distribution-free bound P(unsafe) ≤ ε with probability 1−δ. Certificates hold within 2–5% of empirical violation rates on ToxiGen and AdvBench across four frontier models.

The framework formulates LLM safety certification as a PAC learning problem over a prompt distribution, then computes a sample-efficient bound on the probability of safety predicate violation on future unseen prompts. The bound uses conformal prediction combined with Clopper-Pearson intervals — making it distribution-free and assumption-free about LLM internals. Evaluated on ToxiGen and AdvBench with four frontier models, certificates hold within 2–5% of empirical violation rates at δ = 0.01 with as few as 200 test prompts, and are composable: certificates for sub-policies (refusal gates, output classifiers) can be chained to certify multi-stage safety pipelines.

04
mech-interp preprint

LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

First unlearning testbed with ground-truth parameter-level localization; proves that standard high-metric unlearning does not imply genuine knowledge erasure and quantifies exactly how much better localization could improve robustness to resurfacing attacks.

When an unlearning algorithm achieves high scores on the standard forgetting benchmarks, does it actually remove the knowledge from the weights — or just suppress its expression? LACUNA answers this question by building a testbed where the ground truth is known.

LACUNA: Unlearning Metric vs. Resurfacing Robustness standard unlearning metric (higher = better forget) resurfacing robustness SimNPO & variants (high metric, low robustness) OracleGrad (oracle localization) localization gap
LACUNA exposes the localization gap: standard unlearning algorithms (SimNPO and variants) achieve high unlearning metric scores but low resurfacing robustness because they do not target the responsible weights. OracleGrad — with oracle access to ground-truth parameter localization — achieves both, quantifying the gap that better mechanistic localization methods could close.

LACUNA injects PII of synthetic individuals into predefined weight subsets of 1B and 7B OLMo models via masked continual pretraining, creating a controlled testbed with oracle knowledge of which parameters encode each to-be-forgotten fact. Standard unlearning algorithms achieve high traditional unlearning metrics without targeting the responsible weights, leaving models vulnerable to resurfacing attacks that recover the supposedly forgotten information. OracleGrad — a simple baseline with oracle parameter localization — achieves optimal forget/retain balance and maximal resurfacing robustness, quantifying the gap that better mechanistic localization could close.

05
mech-interp preprint

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

Establishes that SAE family is a more important methodological choice than model scale for causal interventional claims — a concrete warning for safety practitioners selecting SAE families for probing or steering experiments.

If you ablate a feature that reliably activates on a single token type and the model's probability of generating that token doesn't drop significantly, the feature wasn't doing the causal work. Testing 3.9 million such features, this paper finds the answer depends more on which SAE family you trained than on which model you're studying.

Single-Token Feature Causal Necessity (BH-significant, 178/208 conditions) GemmaScope 86% BatchTopK 82% LlamaScope ~40% 0% 86% cross-family gap > within-family scaling effect
Single-token feature causal necessity rates by SAE family. GemmaScope and BatchTopK features are causally anchored (86% and 82% BH-significant); LlamaScope features are locally redundant (~40%). The cross-family gap exceeds the within-family scaling effect — SAE family selection matters more than model scale for interventional validity.

The paper uses single-token SAE features as ground-truth-verifiable diagnostics: ablating a feature activating on token t must significantly reduce log-probability for t to count as causally necessary. Testing 3.9 million features from six models across three SAE families, BH-significant logit reductions appear in 178 of 208 conditions. Single-token features cluster 4.7× tighter in decoder space than all features and concentrate in early layers. The key finding is that cross-family variation exceeds within-family scaling effects, directly implying that SAE family is a first-order methodological choice for anyone relying on SAE interventions for safety work.

06
mech-interp AI security preprint

Sparse Autoencoders are Capable LLM Jailbreak Mitigators

Demonstrates the clearest pragmatic safety payoff for SAE interpretability to date: SAE-guided jailbreak defense outperforms dense activation steering on all four tested models and 12 jailbreak categories, with the largest gap on out-of-distribution attacks.

Dense activation steering for jailbreak defense conflates features active in harmful jailbreak context with features active in all harmful queries — including benign ones. CC-Delta separates them by comparing activations with and without jailbreak context, then steers only on the difference.

SAE activations with jailbreak f_harm + f_jb SAE activations same query, no jailbreak f_harm Δ jailbreak-specific SAE features → steer f_jb only Outperforms dense steering on all 4 models × 12 jailbreak types; largest gains on OOD attacks
CC-Delta identifies jailbreak-specific SAE features by contrasting activations for the same harmful request with and without jailbreak context. Steering on only the delta (f_jb) avoids activating on features shared with benign harm queries, reducing over-refusal while maintaining defense coverage — especially on out-of-distribution attacks not seen during defense calibration.

CC-Delta compares token-level SAE sparse activations for the same harmful request framed with and without jailbreak context. The contrastive delta isolates jailbreak-specific features without contaminating features active in both jailbreak and benign contexts. A mean-shift steering vector applied in SAE latent space at inference time then suppresses only the jailbreak-specific directions. Evaluated across four models and 12 jailbreak categories, CC-Delta outperforms dense activation steering throughout, with the largest gains on out-of-distribution attacks — consistent with the interpretation that SAE decomposition captures jailbreak structure invisible to dense residual-stream methods.

07
mech-interp preprint

Steered LLM Activations are Non-Surjective

Proves a fundamental separation between white-box steerability and black-box prompting that has been assumed but never formally established — with direct implications for how any experiment comparing steering to prompting should be designed and what its results can claim.

Researchers often ask "is there a prompt that produces the same behavior as this steering vector?" This paper proves the answer is almost surely no — the set of states reachable via steering is strictly larger than the set reachable via any prompt.

steered states (vastly larger set) prompt-reachable states (measure zero in steered space) steered state not reachable by prompt ⟹ no prompt reproduces steered behavior · white-box ≠ black-box · decouple evaluations
The set of hidden states reachable via activation steering strictly contains the set reachable via any discrete prompt — with the prompt-reachable manifold having measure zero in the larger space. Consequently, almost surely no prompt can reproduce the same internal behavior as a steering vector, and steering-vs-prompting comparisons cannot assume equivalence.

The paper constructs the manifold of prompt-reachable hidden states and proves it has measure zero in the space of states reachable via activation steering — meaning that for almost any steering vector and magnitude, the induced hidden state lies strictly outside the prompt-reachable manifold. This formal non-surjectivity result establishes a fundamental separation: white-box steerability is not equivalent to black-box prompting, benchmarks that compare the two under equivalence assumptions are methodologically unsound, and evaluations of steering effects should explicitly decouple the white-box and black-box conditions.

08
AI security preprint

MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment

Addresses the key limitation of RL-based multi-turn jailbreaking — each dialogue turn contributes unequally to the final outcome, but prior methods apply a single trajectory-level signal that conflates useful and noise turns — with a principled decomposed credit estimator.

A multi-turn jailbreak attacker needs to learn which conversational turns actually moved the model toward compliance and which were noise — but standard RL broadcasts a single final reward to all turns equally. DC-GRPO gives each turn its own credit signal.

Standard RL (single reward) turn 1 turn 2 turn N ← single R DC-GRPO (per-turn credit) turn 1 r₁+γV₂ turn 2 r₂+γV₃ turn N r_N strong multi-turn ASR cross-model transfer immediate credit (rₜ) + future credit (γVₜ₊₁) decomposed per turn
DC-GRPO replaces a single trajectory-level reward with per-turn credits combining immediate reward rₜ and future-value estimate γVₜ₊₁. Both static- and dynamic-weighted instantiations achieve strong multi-turn attack success rates on black-box frontier models with cross-model transferability.

DC-GRPO (Decomposed Credit Group Relative Policy Optimization) derives a per-turn value function that separates each turn's contribution to the final harmful outcome into immediate and future components, then applies group-relative policy optimization independently per turn. Both static-weighted (fixed γ) and dynamic-weighted (γ learned per dialogue state) instantiations are evaluated on multi-turn jailbreak benchmarks against black-box frontier models, achieving strong attack success rates and cross-model transfer — confirming that per-turn credit decomposition is a broadly applicable improvement over trajectory-level RL for multi-turn adversarial training.

Items 9–11 · Also notable
09
mech-interp preprint

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

Proposes a six-condition reliability filter followed by Borda consensus ranking (F-test, KSG mutual information, Cohen's d) to select SAE features for steering without a learned objective, evaluated on three Gemma-family models. Key methodological finding: raw attribute movement in activation space does not predict generation quality — practitioners need an explicit selection filter before treating SAE steering as effective.


10
AI security preprint

The Balkanization of Execution-Security Research for AI Coding Agents

Systematic review of 39 papers across 17 execution-security categories for AI coding agents; headline finding is 69–98% policy-enforcement failure rates across all categories. Documents 4 patched CVEs and 5 structural gaps, arguing the field lacks the shared ontology and shared benchmarks needed for coordinated progress on agent security.


11
mech-interp preprint

Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs

Attribution and activation patching on Llama-3 models identifies a compact shared set of MLP neurons that is both necessary and sufficient for late-layer arithmetic computation across symbolic arithmetic, natural-language word problems, and Python code. Targeted interventions degrade performance simultaneously in all three modalities, supporting a single surface-form-invariant arithmetic circuit.

Watchlist
← all Research Radar issues · gussand · source