Training-stage durability converged as the week's dominant result from three independent lines. "Everything in Moderation" demonstrated that midtraining data-composition choices create alignment-resistant domain gaps a compensatory alignment pass cannot close; "The Geometry of Refusal" supplied the mechanistic account — post-hoc safety updates sit orthogonally to the capability subspace, masking rather than erasing, so a handful of benign gradient steps restore the masked behavior. ACL 2026's "From Narrow Unlearning to Emergent Misalignment" reinforced this asymmetry: targeted refusal unlearning for a single concept spills across unrelated RAI domains, confirming safety is a distributed, cross-cutting property that cannot be locally excised. Against this backdrop the attack frontier widened: cipher-based jailbreaks moved from a fine-tuning-API threat to an ordinary chat threat; directional ablation was validated on a 320B MoE; and full-duplex speech models opened a new attack surface text-domain alignment never covers. Applied mechanistic interpretability continued its evolution into operational tooling — SAE features driving production security backends, forensic backdoor audits, and dual-direction steering engines.
The most controlled evidence yet that midtraining data composition is a durable decision: every reasoning domain has an interior coverage optimum, and the gaps that composition creates survive subsequent alignment fine-tuning intact — the training stage, not the alignment stage, determines final per-domain performance.
Figure 1: Every KOR-Bench domain shows an interior coverage optimum (fitted peaks 9.9%–35.1%, all within the optimal zone); a compensatory SFT alignment pass raises 116/120 cells but leaves relative per-domain gaps unchanged.
A controlled sweep on Qwen3-8B-Base (replicated at 4B) trains 150 configurations spanning a five-domain reasoning simplex at 5 seeds each. A calibrated permutation test for quadratic interiority gives P ≈ 0.010, confirming interior optima for all five domains in the moderate 10–40% range. The decisive result: a fixed-budget compensatory SFT alignment pass raises 116 of 120 evaluation cells yet leaves relative per-domain gaps essentially unchanged, as does an equal-budget uniform control — confirming that the composition chosen at midtraining persists through downstream alignment.
Post-hoc safety updates land in a suppression regime where the update direction is orthogonal to the capability subspace — a thin, reversible gate over intact capabilities rather than erasure — which is why 100 benign fine-tuning steps collapse refusal in frontier models, and why safety must engage the substrate that forms during pretraining.
Figure: (left) Post-hoc safety updates are orthogonal to the capability subspace — masking, not erasure — and 100 benign gradient steps remove the gate. (right) OLMo-2-1B 267-checkpoint sweep: the safety substrate emerges sharply between ~6B and ~60B pretraining tokens, not gradually.
A kernel-immobility lemma formalizes why an update orthogonal to the capability span can only mask, never erase: the capability directions are geometrically immovable under such updates. One hundred benign fine-tuning steps collapse refusal in both Qwen-2.5-7B and Llama-3-8B Instruct, confirming the suppression-regime account. A 267-checkpoint sweep of OLMo-2-1B then traces the formation of the safety substrate, finding a sharp emergence between ~6B and ~60B pretraining tokens rather than gradual accumulation — the durability gap is a pretraining property, not a post-training artifact, and building robust safety requires engaging the substrate.
The same behavioral refusal metric hides fundamentally different underlying circuits: SFT, reasoning-augmented SFT, and ORPO all produce models that refuse at similar rates, yet their refusal circuits have different topologies, steerabilities, and fragilities. No current method achieves all three desired properties simultaneously.
Figure 1: Three post-training methods measured on three desiderata across Llama-3.1-8B, Gemma-2-9B, and Qwen3-8B. Reasoning-augmented training distributes refusal more broadly (not fragile, steerable) but at the cost of minor capability regression; SFT produces a compact, ablatable circuit; no method achieves all three simultaneously.
Attention head attribution and activation patching across three model families show SFT concentrates refusal in a small, easily ablated head set; reasoning-augmented training distributes it more broadly with a distinct computational signature consistent across all three base models; ORPO is intermediate. Architecture also exerts an independent effect: the same training method installs differently steerable circuits depending on the base model. No method achieves all three simultaneously: (1) refusal not in fragile components, (2) no capability regression on benchmarks, and (3) refusal correctable by small targeted interventions. The finding points toward interpretability-informed training design as necessary for durable, controllable refusal.
Figure: Cipher attack requires only ordinary chat access — no gradient, no fine-tuning API. The model learns the cipher in-context and responds to the harmful exchange without triggering safety behaviors.
Demonstrates that frontier models (Anthropic, Google, OpenAI) acquire arbitrary ciphers from in-context examples alone; harmful exchanges conducted in the learned cipher bypass alignment that holds in plaintext. This moves cipher attacks from a fine-tuning-API threat model to an ordinary chat threat model with no technical barrier to entry.
Figure: Refusal removal by writer type in a 320B MoE. No single writer type carries the full refusal signal; only joint editing of all three reaches 77.6% removal — with no gradient-based training.
Extends directional ablation to GLM-5.3-Flash (320B parameters, 288 routed experts, block-FP8). Editing attention, dense, and routed-expert writer matrices individually removes 3.9%, 1.6%, and 14.8% of refusal respectively; joint editing reaches 77.6% with only a few hundred contrastive prompts — validating the attack at true frontier scale for the first time.
Figure: Single-direction refusal steering (dashed) leaves harmful-continuation features active on wrapped prompts; REINS (solid) simultaneously suppresses harm features and amplifies refusal features, recovering genuine refusals without output degeneration.
On GUISE (harmful requests hidden in complex wrappers), existing single-direction SAE steering fails because harmful-continuation features stay active; REINS inhibits them simultaneously with refusal-feature amplification. Harmful-response rate drops markedly and genuine refusals increase, while capability benchmarks are largely preserved — in contrast to baselines that either fail to refuse or degenerate outputs.
Figure: Unlearning refusal for a single RAI concept depresses refusal scores across all seven domains tested (bars all negative); unlearning the Safety concept causes the broadest collateral damage.
ACL 2026 Short Paper (Amazon). Unlearning refusal for one RAI concept (Cybersecurity or Safety) on Mistral-7B-v0.3 and Qwen2.5-7B induces emergent misalignment well outside the targeted domain; Safety-concept unlearning causes the broadest spillover, depressing refusal in domains such as bias, toxicity, and privacy. Peer-reviewed evidence that safety is a cross-cutting, distributed property that cannot be locally excised.
Figure: A jailbreak vector extracted from Class A (training source) transfers to suppress semantically unrelated Classes B and C — supporting a shared harmfulness-suppression mechanism and enabling Latent Guard.
EACL 2026 Long Paper (LMU / Anthropic). Effective jailbreaks measurably lower the model's internal harmfulness representation; a jailbreak vector extracted from one semantic class suppresses jailbreaks from other classes. Latent Guard, built from this shared mechanism, matches Llama Guard 3 8B performance while reducing over-refusals, and is robust to fine-tuning attacks. Evaluated on Vicuna 13B/7B v1.5, Qwen1.5 14B Chat, and MPT 7B Chat.
Figure: EraseSAE pipeline for text-to-video models: Partitioned Convolutional SAE decomposes spatiotemporal activations; contrastive attribution isolates concept-specific feature kernels; timestep masks confine erasure to only where the target concept is active.
ECCV 2026. Extends SAE-based concept erasure to DiT text-to-video models (HunyuanVideo, CogVideoX-5b), achieving precise celebrity-identity and nudity erasure with minimal quality degradation. Key advance over coarse-grained methods: monosemantic feature-level operation prevents collateral erasure of adjacent concepts. The Partitioned Convolutional SAE design is the first architecture handling the spatiotemporal structure of video diffusion activations.
Figure: Trigger detection F1 (blue) is high and stable across many layers; causal control of the backdoor behavior (red) is concentrated in a narrower set of layers. Ablating the best detectors does not reliably remove the backdoor.
Uses SAEs on 1B and 8B models to forensically trace a language-switch backdoor (trigger → output in French/German). SAE features detect triggered prompts with near-perfect F1, but the features that detect the trigger are not the features that control the behavior — ablating detectors does not reliably remove the backdoor. Any SAE-based backdoor audit needs distinct feature sets for detection versus removal.
Figure: DuplexJail delivers spoken interruption after the harmful request ends, overriding the model's nascent refusal in the streaming audio channel; AdvBench ASR jumps from ~7% to 40–49% on PersonaPlex variants.
Full-duplex speech models process user audio concurrently with output generation, creating an attack surface text-domain safety alignment never encounters. Fixed-delay spoken interruption raises whole-response AdvBench attack success rate by +33.8 and +39.3 pp (to 40.3% and 48.7%) on PersonaPlex and PersonaPlex-RL, representing an entirely new threat model for voice-capable frontier systems.
Figure: Safety-relevant and language-identity SAE features overlap across layers and languages (darker = more overlap); the causal test confirms ablating safety features shifts both harmful-compliance rate and output language simultaneously.
Applies residual-stream SAEs at every layer of Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Gemma-2-9B-IT across eight languages; safety-relevant features are geometrically entangled with language-identity features in architecture-specific layer bands. Ablating a language's safety features raises harmful-compliance rates and shifts output language — a purely safety-targeted feature-level intervention is not available, with direct implications for why English-trained safety alignment transfers unevenly across languages.
Figure: Truth and prescribed-action probe AUROCs sum to exactly 1.0 across all 751 tested cell-layer pairs — the two hypotheses are perfectly aliased and unidentifiable from compliant-context labels alone.
A fundamental negative result for deception probes: when compliant behavior and truthful reporting coincide, a truth probe and a prescribed-action probe solve identical optimization problems on compliant contexts, and on rival contexts their AUROCs are exact complements (sum = 1.0 to floating-point precision across 751 cell-layer pairs). Disambiguation requires randomized codebooks separating output symbol from semantic action and fitting on mixed compliant + rival contexts — a necessary methodological condition for any deception detector.
Watchlist for next week
Locating and Steering Refusal Beyond Attention — Cross-architecture refusal transfer: rigid rotation aligns transformer and SSM residual streams; a probe from a transformer flags SSM harmful inputs without re-training; ablating the aligned direction from an SSM bypasses its refusal. Watch for systematic cross-architecture evaluation at scale.
Feature Drift: How Fine-Tuning Repurposes Representations in LLMs (EACL 2026 Findings — ACL Anthology only, no arXiv ID) — Base SAEs applied to chat-model activations degrade 5–10% vs. chat-trained SAEs; feature space survives post-training but activation distribution shifts. Critical for whether base-model SAE safety tooling holds after alignment tuning. Watch for an arXiv posting.
Representational Alignment Yields Generalizable Safety in Language Models — RSO directly aligns LLM latent geometry with human moral judgment prototypes; improved generalization to adversarial harm rephrasings across 23 models. Watch for scale-up and red-team evaluation against adaptive rephrasings.
ObserverBench: Testing Mechanistic Estimates for Intervention and Control — Benchmark decoupling estimation accuracy from action-choice quality for interp-guided interventions; reveals a substantial accuracy–control gap in circuit-intervention tasks. Watch for extension to safety-relevant steering decisions at deployment scale.
Notes: No daily reports exist for September 7, 9, 12, or 13; this weekly draws from the Sep 8, Sep 10, and Sep 11 dailies. Papers from September 12–13 may be missing from this edition.
Feature Drift (EACL 2026 Findings) has no arXiv ID — it exists only at an ACL Anthology URL. The radar_index.py ledger cannot track anthology-only papers; this entry is placed in the Watchlist to avoid a recurring blind spot.
All 13 ranked items verified against reports/index/seen.tsv; weeklies are checked against earlier weeklies only (aggregating the week's dailies is a weekly's job).