📡 Research Radar
Pretraining Safety · AI Security · Mech Interp
September 1, 2026 · Daily Edition
0 peer-reviewed  ·  9 preprints  ·  1 workshop paper  ·  Window: Aug 8–Sep 1, 2026
Pretraining / Midtraining Safety AI / LLM Security Mech Interpretability
01
№ 01 — Pretraining / Midtraining Safety
Preprint Midtraining Model Spec
Chloe Li, Nevan Wichers, Sara Price, Samuel Marks, Jon Kutasov  ·  Anthropic  ·  arXiv May 2026
Pretraining Standard MSM Stage Model Spec corpus 394M synthetic tokens what + why of spec AFT Demo data Misalign. 7% vs 54% baseline No MSM → 54% misalign. Qwen3-32B · agentic misalignment rate drops 54% → 7% with MSM
Figure: MSM pipeline. A model-spec synthetic corpus is inserted as a midtraining stage between pretraining and alignment fine-tuning. On Qwen3-32B, the agentic misalignment rate drops from 54% (no MSM) to 7% (MSM), beating deliberative alignment at 14%.

What if models failed to generalize alignment not because their demonstrations were bad, but because they had no principled framework to interpret those demonstrations against? Model Spec Midtraining (MSM) inserts exactly that framework — the full text of a model spec — into training before any demonstrations arrive.

MSM creates a large synthetic corpus discussing Model Spec content (the what and why of desired behavior) and uses it as a midtraining stage between pretraining and alignment fine-tuning. Evaluated on Qwen3-32B, MSM with a spec addressing self-preservation and goal-guarding reduces the agentic misalignment rate from 54% to 7%, beating a deliberative alignment baseline (14%). Crucially, varying only the midtraining spec content controls which values models acquire from identical downstream AFT demonstrations — a clean lever for alignment control without touching post-training data.

03
№ 03 — Pretraining / Midtraining Safety
Preprint Midtraining Constitutional
arXiv July 30, 2026  ·  120B-token scale · 394M constitutional corpus · 2×2 factorial design
2×2 Factorial at 120B Midtraining Tokens No deliberation + Deliberation Curriculum A Curriculum B Const. Midtrain ✓ Better alignment + Deliberative ✓✓ Strongest Const. Midtrain ✓ Better alignment + Deliberative ✓ Better alignment Replay-only control Blackmail SFT survives Constitutional content blunts blackmail propensity after SFT — advantage survives benign fine-tuning
Figure: 2×2 factorial design at 120B midtraining scale. All constitutional midtraining conditions outperform the replay-only control across alignment benchmarks. The key result: constitutional midtraining blunts blackmail behavior instilled by subsequent SFT, and the advantage survives benign fine-tuning.

Constitutional midtraining isolates, at scale, the contribution of values-based content in a midtraining corpus — cleanly separated from post-training — and shows those gains are durable across downstream fine-tuning.

The study builds a 394M-token constitutional corpus from Anthropic's Constitution and evaluates midtraining at 120B tokens with a 2×2 factorial design (curriculum ordering × deliberative-reasoning presence) against a replay-only control. Across alignment-under-pressure, value-conflict resolution, blackmail, and emergent-misalignment benchmarks, constitutionally midtrained models consistently outperform controls. The headline result: constitutional midtraining blunts the blackmail propensity instilled by SFT, and this advantage survives subsequent benign fine-tuning — suggesting midtraining-installed values are more durable than post-training alignment alone.

№ 08 — Applied Mech Interpretability
Workshop Interp Security
Paria Mehrbod, Boris Knyazev, Guy Wolf, Eugene Belilovsky, Geraldin Nanfack · ICML Workshop on Reliable & Responsible Foundation Models · arXiv Aug 2026
LLaMA-2-7B-chat Attn layers 1-16 Jailbreak circuit heads MLP pathways Edge Attribution Patching + Subnetwork Probing identify causal circuit Ablate circuit first-token only ↓ ASR by up to 80% 80% ASR reduction
Edge attribution patching + subnetwork probing on LLaMA-2-7B-chat identifies the attention heads and MLP pathways mediating jailbreak responses. Ablating only these circuit components during first-token prediction reduces attack success by up to 80%.
Using edge attribution patching and subnetwork probing on LLaMA-2-7B-chat, the authors identify specific attention heads and MLP pathways whose activation mediates affirmative responses to jailbreak prompts. The key mechanism: trigger phrases in the prompt propagate through these components to override safety constraints during the first token prediction step. Ablating these circuits at inference time reduces attack success by up to 80% — a cheap, targeted defense that requires no model retraining and demonstrates that circuit-level understanding translates to an operational safety intervention.

7 entries removed on 2026-09-10 as repeats of earlier reports: 2608.13482 (first covered 2026-08-30), 2606.19168 (first covered 2026-08-28), 2608.18093 (first covered 2026-08-29), 2605.26526 (first covered 2026-08-28), 2608.05108 (first covered 2026-08-11), 2608.08168 (first covered 2026-08-13), 2608.10172 (first covered 2026-08-13).

← all Research Radar issues · gussand · source