Research Radar
Week 38 · September 7–13, 2026
4 peer-reviewed · 9 preprints · 0 forum/blog  ·  backfilled 2026-09-29
AI Safety · Alignment · Mech Interp
Theme of the Week

Training-stage durability converged as the week's dominant result from three independent lines. "Everything in Moderation" demonstrated that midtraining data-composition choices create alignment-resistant domain gaps a compensatory alignment pass cannot close; "The Geometry of Refusal" supplied the mechanistic account — post-hoc safety updates sit orthogonally to the capability subspace, masking rather than erasing, so a handful of benign gradient steps restore the masked behavior. ACL 2026's "From Narrow Unlearning to Emergent Misalignment" reinforced this asymmetry: targeted refusal unlearning for a single concept spills across unrelated RAI domains, confirming safety is a distributed, cross-cutting property that cannot be locally excised. Against this backdrop the attack frontier widened: cipher-based jailbreaks moved from a fine-tuning-API threat to an ordinary chat threat; directional ablation was validated on a 320B MoE; and full-duplex speech models opened a new attack surface text-domain alignment never covers. Applied mechanistic interpretability continued its evolution into operational tooling — SAE features driving production security backends, forensic backdoor audits, and dual-direction steering engines.

01
alignment midtraining preprint

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

The most controlled evidence yet that midtraining data composition is a durable decision: every reasoning domain has an interior coverage optimum, and the gaps that composition creates survive subsequent alignment fine-tuning intact — the training stage, not the alignment stage, determines final per-domain performance.

Domain coverage share (%) Task accuracy 0 20 40 60 80 optimal zone 10–40% Domain A Domain B Domain C Domain D Domain E
Figure 1: Every KOR-Bench domain shows an interior coverage optimum (fitted peaks 9.9%–35.1%, all within the optimal zone); a compensatory SFT alignment pass raises 116/120 cells but leaves relative per-domain gaps unchanged.

A controlled sweep on Qwen3-8B-Base (replicated at 4B) trains 150 configurations spanning a five-domain reasoning simplex at 5 seeds each. A calibrated permutation test for quadratic interiority gives P ≈ 0.010, confirming interior optima for all five domains in the moderate 10–40% range. The decisive result: a fixed-budget compensatory SFT alignment pass raises 116 of 120 evaluation cells yet leaves relative per-domain gaps essentially unchanged, as does an equal-budget uniform control — confirming that the composition chosen at midtraining persists through downstream alignment.


02
alignment mech-interp preprint

The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists

Post-hoc safety updates land in a suppression regime where the update direction is orthogonal to the capability subspace — a thin, reversible gate over intact capabilities rather than erasure — which is why 100 benign fine-tuning steps collapse refusal in frontier models, and why safety must engage the substrate that forms during pretraining.

Suppression Regime capability subspace safety update (orthogonal) 100 benign gradient steps erase the gate → refusal collapses OLMo-2-1B Checkpoint Sweep 1B 6B 60B 300B pretraining tokens safety substrate sharp emergence
Figure: (left) Post-hoc safety updates are orthogonal to the capability subspace — masking, not erasure — and 100 benign gradient steps remove the gate. (right) OLMo-2-1B 267-checkpoint sweep: the safety substrate emerges sharply between ~6B and ~60B pretraining tokens, not gradually.

A kernel-immobility lemma formalizes why an update orthogonal to the capability span can only mask, never erase: the capability directions are geometrically immovable under such updates. One hundred benign fine-tuning steps collapse refusal in both Qwen-2.5-7B and Llama-3-8B Instruct, confirming the suppression-regime account. A 267-checkpoint sweep of OLMo-2-1B then traces the formation of the safety substrate, finding a sharp emergence between ~6B and ~60B pretraining tokens rather than gradual accumulation — the durability gap is a pretraining property, not a post-training artifact, and building robust safety requires engaging the substrate.


03
EMNLP 2026 mech-interp alignment

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

The same behavioral refusal metric hides fundamentally different underlying circuits: SFT, reasoning-augmented SFT, and ORPO all produce models that refuse at similar rates, yet their refusal circuits have different topologies, steerabilities, and fragilities. No current method achieves all three desired properties simultaneously.

SFT Reasoning-Aug ORPO Not fragile No cap. regress. Steerable ✗ ✓ ✗ ✓ ✗ ✓ ~ ✓ ~ No single method satisfies all three properties simultaneously
Figure 1: Three post-training methods measured on three desiderata across Llama-3.1-8B, Gemma-2-9B, and Qwen3-8B. Reasoning-augmented training distributes refusal more broadly (not fragile, steerable) but at the cost of minor capability regression; SFT produces a compact, ablatable circuit; no method achieves all three simultaneously.

Attention head attribution and activation patching across three model families show SFT concentrates refusal in a small, easily ablated head set; reasoning-augmented training distributes it more broadly with a distinct computational signature consistent across all three base models; ORPO is intermediate. Architecture also exerts an independent effect: the same training method installs differently steerable circuits depending on the base model. No method achieves all three simultaneously: (1) refusal not in fragile components, (2) no capability regression on benchmarks, and (3) refusal correctable by small targeted interventions. The finding points toward interpretability-informed training design as necessary for durable, controllable refusal.


Items 4–13 · Also notable
04
AI security jailbreak preprint

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning

Stage 1 Teach cipher in-context Stage 2 Send harmful request in cipher Stage 3 Model responds in cipher — no refusal no API access no fine-tuning Alignment holds in plaintext; safety bypassed in cipher representation
Figure: Cipher attack requires only ordinary chat access — no gradient, no fine-tuning API. The model learns the cipher in-context and responds to the harmful exchange without triggering safety behaviors.

Demonstrates that frontier models (Anthropic, Google, OpenAI) acquire arbitrary ciphers from in-context examples alone; harmful exchanges conducted in the learned cipher bypass alignment that holds in plaintext. This moves cipher attacks from a fine-tuning-API threat model to an ordinary chat threat model with no technical barrier to entry.


05
AI security mech-interp preprint

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

3.9% Attention 1.6% Dense 14.8% Routed Exp. 77.6% Joint 0% 40% 80% Refusal removal rate — GLM-5.3-Flash (320B MoE)
Figure: Refusal removal by writer type in a 320B MoE. No single writer type carries the full refusal signal; only joint editing of all three reaches 77.6% removal — with no gradient-based training.

Extends directional ablation to GLM-5.3-Flash (320B parameters, 288 routed experts, block-FP8). Editing attention, dense, and routed-expert writer matrices individually removes 3.9%, 1.6%, and 14.8% of refusal respectively; joint editing reaches 77.6% with only a few hundred contrastive prompts — validating the attack at true frontier scale for the first time.


06
mech-interp steering preprint

REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features

Wrapped harmful prompt (role-play, task) harm features still active REINS steers SAE feature space ↓ suppress harm features ↑ amplify refusal features two directions simultaneously baseline: harmful output REINS: genuine refusal
Figure: Single-direction refusal steering (dashed) leaves harmful-continuation features active on wrapped prompts; REINS (solid) simultaneously suppresses harm features and amplifies refusal features, recovering genuine refusals without output degeneration.

On GUISE (harmful requests hidden in complex wrappers), existing single-direction SAE steering fails because harmful-continuation features stay active; REINS inhibits them simultaneously with refusal-feature amplification. Harmful-response rate drops markedly and genuine refusals increase, while capability benchmarks are largely preserved — in contrast to baselines that either fail to refuse or degenerate outputs.


07
ACL 2026 alignment unlearning

From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs

cyber safety toxic. bias sens. med/legal privacy 0 − RAI domain refusal score change (unlearning one concept)
Figure: Unlearning refusal for a single RAI concept depresses refusal scores across all seven domains tested (bars all negative); unlearning the Safety concept causes the broadest collateral damage.

ACL 2026 Short Paper (Amazon). Unlearning refusal for one RAI concept (Cybersecurity or Safety) on Mistral-7B-v0.3 and Qwen2.5-7B induces emergent misalignment well outside the targeted domain; Safety-concept unlearning causes the broadest spillover, depressing refusal in domains such as bias, toxicity, and privacy. Peer-reviewed evidence that safety is a cross-cutting, distributed property that cannot be locally excised.


08
EACL 2026 AI security mech-interp

Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

latent space jailbreak vector ↓ harm Class A (train source) Class B (unseen) Class C (unseen) → Latent Guard: repurposes the vector as a safeguard (≈ Llama Guard 3 8B)
Figure: A jailbreak vector extracted from Class A (training source) transfers to suppress semantically unrelated Classes B and C — supporting a shared harmfulness-suppression mechanism and enabling Latent Guard.

EACL 2026 Long Paper (LMU / Anthropic). Effective jailbreaks measurably lower the model's internal harmfulness representation; a jailbreak vector extracted from one semantic class suppresses jailbreaks from other classes. Latent Guard, built from this shared mechanism, matches Llama Guard 3 8B performance while reducing over-refusals, and is robust to fine-tuning attacks. Evaluated on Vicuna 13B/7B v1.5, Qwen1.5 14B Chat, and MPT 7B Chat.


09
ECCV 2026 mech-interp concept-erasure

EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders

1. Decompose Part. Conv. SAE spatiotemporal feats 2. Attribute Contrastive attribution isolate concept kernels 3. Erase Timestep-resolved masks → surgical DiT text-to-video model (HunyuanVideo / CogVideoX-5b) Precise concept erasure · minimal quality degradation
Figure: EraseSAE pipeline for text-to-video models: Partitioned Convolutional SAE decomposes spatiotemporal activations; contrastive attribution isolates concept-specific feature kernels; timestep masks confine erasure to only where the target concept is active.

ECCV 2026. Extends SAE-based concept erasure to DiT text-to-video models (HunyuanVideo, CogVideoX-5b), achieving precise celebrity-identity and nudity erasure with minimal quality degradation. Key advance over coarse-grained methods: monosemantic feature-level operation prevents collateral erasure of adjacent concepts. The Partitioned Convolutional SAE design is the first architecture handling the spatiotemporal structure of video diffusion activations.


10
mech-interp AI security backdoor preprint

LLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders

Detection F1 (high, stable) Causal control (concentrated, narrow) Layer depth → Score Detection ≠ control: different layers, different feature sets
Figure: Trigger detection F1 (blue) is high and stable across many layers; causal control of the backdoor behavior (red) is concentrated in a narrower set of layers. Ablating the best detectors does not reliably remove the backdoor.

Uses SAEs on 1B and 8B models to forensically trace a language-switch backdoor (trigger → output in French/German). SAE features detect triggered prompts with near-perfect F1, but the features that detect the trigger are not the features that control the behavior — ablating detectors does not reliably remove the backdoor. Any SAE-based backdoor audit needs distinct feature sets for detection versus removal.


11
AI security speech preprint

DuplexJail: Safety Alignment Breaks Under Spoken Interruption in Full-Duplex Models

model generation (streaming text + audio output) User audio in harmful request nascent refusal spoken interruption (fixed delay) harmful output resumed AdvBench ASR: baseline ~7% → 40–49% (+33–39 pp)
Figure: DuplexJail delivers spoken interruption after the harmful request ends, overriding the model's nascent refusal in the streaming audio channel; AdvBench ASR jumps from ~7% to 40–49% on PersonaPlex variants.

Full-duplex speech models process user audio concurrently with output generation, creating an attack surface text-domain safety alignment never encounters. Fixed-delay spoken interruption raises whole-response AdvBench attack success rate by +33.8 and +39.3 pp (to 40.3% and 48.7%) on PersonaPlex and PersonaPlex-RL, representing an entirely new threat model for voice-capable frontier systems.


12
mech-interp alignment multilingual preprint

When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs

Safety ↔ Language-Identity Feature Overlap Early layers Mid layers Late layers EN ZH AR HI Ablating safety features also shifts output language — entanglement is causal
Figure: Safety-relevant and language-identity SAE features overlap across layers and languages (darker = more overlap); the causal test confirms ablating safety features shifts both harmful-compliance rate and output language simultaneously.

Applies residual-stream SAEs at every layer of Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Gemma-2-9B-IT across eight languages; safety-relevant features are geometrically entangled with language-identity features in architecture-specific layer bands. Ablating a language's safety features raises harmful-compliance rates and shifts output language — a purely safety-targeted feature-level intervention is not available, with direct implications for why English-trained safety alignment transfers unevenly across languages.


13
mech-interp AI-control preprint

The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

Truth probe AUROC Prescribed-action AUROC 0.5 1.0 1.0 AUROC(truth) + AUROC(action) = 1.0 exact, to floating-point precision (751 cell-layer pairs)
Figure: Truth and prescribed-action probe AUROCs sum to exactly 1.0 across all 751 tested cell-layer pairs — the two hypotheses are perfectly aliased and unidentifiable from compliant-context labels alone.

A fundamental negative result for deception probes: when compliant behavior and truthful reporting coincide, a truth probe and a prescribed-action probe solve identical optimization problems on compliant contexts, and on rival contexts their AUROCs are exact complements (sum = 1.0 to floating-point precision across 751 cell-layer pairs). Disambiguation requires randomized codebooks separating output symbol from semantic action and fitting on mixed compliant + rival contexts — a necessary methodological condition for any deception detector.


Watchlist for next week

Notes: No daily reports exist for September 7, 9, 12, or 13; this weekly draws from the Sep 8, Sep 10, and Sep 11 dailies. Papers from September 12–13 may be missing from this edition.

Feature Drift (EACL 2026 Findings) has no arXiv ID — it exists only at an ACL Anthology URL. The radar_index.py ledger cannot track anthology-only papers; this entry is placed in the Watchlist to avoid a recurring blind spot.

All 13 ranked items verified against reports/index/seen.tsv; weeklies are checked against earlier weeklies only (aggregating the week's dailies is a weekly's job).

← all Research Radar issues · gussand · source