RESEARCH RADAR
Week 37 · Aug 8 – Sep 7, 2026
0 peer-reviewed · 12 preprints · 1 lab blog
Pretraining Safety · AI Security · Mech Interp
Theme of the week

Pretraining and midtraining safety interventions have converged into a coherent research agenda, with five independent groups each intervening at a different point in the pipeline: filtering the corpus for dangerous capabilities (Deep Ignorance), injecting safety reflections into the raw pretraining stream (SRP), installing an aligned persona from token zero (Synthetic Persona Pretraining), training on specification documents before alignment fine-tuning (Model Spec Midtraining), and grounding preference-pair synthesis in a formal model spec (SpecAlign). A fresh mechanistic result this week (When Safety Routing Breaks, 2609.01455) explains why post-training alignment degrades so readily under benign fine-tuning: the safety direction is low-rank and concentrated in output-side MLP modules, making it selectively vulnerable to re-sharpening with only ~100 gradient steps. On the security side, abliteration matured from attack to subfield: two August papers study refusal vector geometry and propose dedicated defenses (AMRA, ART) that target the root cause — extractability of the refusal direction — rather than its robustness after extraction.

01
pretraining-safety mech-interp preprint

When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

Safety alignment breaks catastrophically after just ~100 benign fine-tuning examples — a reproducible empirical fact without a mechanistic account, until now. This paper finds that safety is not distributed evenly through the network but routed through a low-rank pathway concentrated in output-side MLP modules, and benign gradients preferentially re-sharpen exactly that pathway.

Fisher magnitude Eigenvector rank (sorted) General capability Fisher Safety Fisher (low-rank) Output-side MLP modules High Low Re-sharpened by 100 benign steps
Safety Fisher spectrum is low-rank relative to general-capability Fisher; output-side MLP modules concentrate the safety pathway, which is selectively re-sharpened by ~100 benign fine-tuning gradient steps — collapsing safety routing while leaving general utility intact.

The paper decomposes parameter importance via a Fisher information analysis, finding the safety Fisher matrix is low-rank compared to general capability Fisher, and alignment training flattens safety geometry while preserving a tight output-routing pathway. Benign gradient updates selectively re-sharpen this pathway in output-side MLP modules, flipping the routing without touching general utility geometry. This explains three puzzles: why safety collapses after few benign steps, why general utility is largely unaffected, and why a handful of safety examples can restore refusal (the underlying safety representation is preserved; only the routing output weights need resetting). LoRA and ASAM mitigate early collapse by suppressing output-side sharpness, but both protections erode at scale (5000+ samples).


02
pretraining-safety preprint

Synthetic Persona Pretraining: Alignment from Token Zero

The model learns what kind of entity it is during pretraining — before any RLHF. Synthetic Persona Pretraining asks whether explicitly installing an aligned assistant persona at pretraining time, via first-person moral reflections annotated on documents, produces more durable and robust alignment than leaving persona to emerge from post-training alone.

SPP Baseline Persona-annotated pretraining (500B tok) Persona binding Strong constitution adherence ✓ Standard pretraining Weaker adherence no value priority shift Advantage scales with pretraining budget; late-stage SPP insertion is substantially weaker. SPP: ↑ jailbreak robustness · ↓ OOD misalignment rate vs. standard pretraining + RLHF
SPP annotates pretraining documents with first-person moral reflections drawn from a constitution, establishing the assistant persona from token zero. Early intervention is substantially stronger than introducing SPP late in pretraining — the alignment benefit scales with pretraining budget.

SPP works in three stages: annotate pretraining documents with first-person moral reflections derived from a constitution; pretrain on this annotated corpus so the model learns the aligned persona as part of its base distribution; bind the pretraining persona to the assistant role in post-training. Experiments scaling to 3B parameters / 500B tokens show SPP improves constitution following and jailbreak robustness, and reduces misalignment rate in out-of-distribution moral dilemmas. Introducing SPP only at the end of pretraining yields substantially weaker results — confirming that early intervention, not just more aligned data, drives the gain. Connects directly to the MSM line (#03 below) and the alignment-pretraining discourse literature.


03
midtraining-safety alignment preprint Anthropic lab blog

Model Spec Midtraining: Improving How Alignment Training Generalizes

Alignment fine-tuning can fail to generalize because demonstrations underspecify the intended behavior — especially for complex principles like corrigibility. MSM (Model Spec Midtraining) shapes the model's conceptual vocabulary for alignment before fine-tuning by training it on documents that discuss the model spec, so the subsequent RLHF data can generalize far more broadly.

Agentic Misalignment Rate — Qwen3-32B 54% Baseline 14% Deliberative alignment 7% MSM MSM achieves 2× alignment performance with ~10% of midtraining data
MSM reduces agentic misalignment rate from 54% → 7% on Qwen3-32B (beating deliberative alignment at 14%) while using only ~10% of the midtraining data for 2× alignment performance. Result from Anthropic researchers; also posted at alignment.anthropic.com.

MSM trains models on synthetic documents that discuss, explain, and reason about the model specification — between pretraining and alignment fine-tuning. By building a conceptual vocabulary for alignment concepts before seeing demonstration data, subsequent RLHF generalizes far more broadly. On Qwen3-32B with a self-preservation / goal-guarding spec, agentic misalignment drops from 54% → 7%, outperforming deliberative alignment (14%). The method is 2× more alignment-efficient using only ~10% of the midtraining data, driven by out-of-distribution generalization. Anthropic released this alongside the alignment.anthropic.com page for the Model Spec, suggesting this may inform production training pipelines.


04
pretraining-safety preprint

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection

Data filtering makes pretraining data safe — but a model can still compose benign knowledge and capabilities into unsafe behaviors. Safety Reflection Pretraining (SRP) inserts periodic self-monitoring reflections into the pretraining corpus so the model learns to reason about safety as part of language modeling, not just as an output-suppression rule.

Safety Reflection Pretraining Pipeline Web doc ⟨safety reflection⟩ "I should not..." Web doc ⟨safety reflection⟩ "This could be used..." SRP (1.7B, FineWeb-Edu) vs. Baselines Inference ASR Filter-only SRP ✓ SRP reduces inference-stage + fine-tuning attack success rates Comparable general benchmark performance to baseline MedSafetyWorld ablations confirm advantage over filtering + rewriting
SRP inserts periodic safety reflection annotations into the pretraining stream. Experiments on 1.7B models (FineWeb-Edu) show SRP substantially reduces inference-stage and fine-tuning attack success rates, while matching baseline on general benchmarks.

SRP pretrains 1.7B-parameter models on FineWeb-Edu augmented with periodic safety reflections — synthetic annotations that elicit self-monitoring about whether current context is eliciting harmful outputs. Compared with data filtering and data rewriting baselines, SRP substantially reduces success rates of inference-stage jailbreaks and fine-tuning attacks while maintaining comparable performance on general understanding and instruction-following benchmarks. The paper introduces MedSafetyWorld, a fully controlled synthetic environment that measures whether safety generalizes from safe data composition; ablations confirm SRP's advantage: the model learns to apply safety reasoning in novel compositions, whereas filtering-only models inherit no such inductive bias.


05
pretraining-safety AI security preprint / under review

Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs

Post-training alignment is famously brittle to adversarial fine-tuning. Deep Ignorance tests whether the right counterfactual is "remove the capability, not just the behavior" — filtering dual-use content from pretraining so the model never acquires the capability in the first place, making it substantially harder to restore than an RLHF-removed behavior.

Biorisk knowledge accuracy vs. adversarial fine-tuning Fine-tuning steps (biothreat text) Accuracy 0 100% chance Baseline (post-training) Deep Ignorance (filtered) 0 2.5K 5K 7.5K 10K
Filtered 6.9B models (Deep Ignorance) maintain near-random biorisk knowledge accuracy across 10,000 adversarial fine-tuning steps — over an order of magnitude more tamper-resistant than post-training alignment baselines (EleutherAI / UK AI Security Institute / Oxford).

The paper introduces a multi-stage scalable pipeline for filtering dual-use content from pretraining corpora: a classifier identifies biothreat-relevant documents, followed by targeted removal. 6.9B models pretrained on the filtered corpus perform at near-random-chance on biorisk knowledge evaluations and resist adversarial fine-tuning for more than 300 million tokens of biothreat material — over an order of magnitude more tamper-resistant than post-training alignment baselines. Regressions on general benchmarks are minimal (limited to closely related biology knowledge). The key asymmetry: reinstalling a capability never learned requires substantially more data than removing a behavior that was explicitly trained, making pretraining filtering a structural rather than behavioral safeguard.


06
midtraining-safety alignment preprint

SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data

Frontier labs increasingly need models to conform to detailed provider-authored specs, not just aggregate human preference. SpecAlign synthesizes fine-grained alignment preference data directly from spec documents, combining structured rule annotation, controllable example generation, and multi-agent adversarial synthesis to produce boundary-aware preference pairs.

Model Spec document Rule annotation (fine-grained) Controllable instantiation (compliant/violating) Multi-agent adversarial synthesis Boundary- aware prefs ✓ Consistently improves spec-rule compliance vs. RLHF/DPO baselines across multiple backbone models while preserving general capabilities and avoiding over-conservative refusal Key advantage: boundary-aware pairs capture where compliance ends, not just positive examples
SpecAlign combines structured rule annotation (parsing specs into fine-grained atomic rules), controllable instantiation (compliant and violating examples per rule), and multi-agent adversarial synthesis to produce boundary-aware preference pairs for LLM alignment.

SpecAlign decomposes the alignment problem as: parse the spec into atomic rules with explicit compliance conditions; generate compliant and boundary-violating completions for each rule; and deploy a red-team agent to probe boundary cases, refining preference pairs. Experiments across multiple backbone models and independent model specifications show consistent improvement in rule compliance vs. RLHF and DPO baselines on held-out spec evaluation, while preserving general capability scores and not amplifying over-conservative refusal. The key advantage over prior spec-following approaches: SpecAlign produces boundary-aware pairs that capture exactly where compliance ends, enabling models to learn the intended decision boundary rather than just positive examples.


07
AI security pretraining-safety preprint

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

Alignment-preserving fine-tuning defenses like TAR and SEAM are typically benchmarked against expensive gradient-based adversarial fine-tuning. Two cheap gradient-free alternatives — abliteration and prefilling — turn out to be sufficient to break them, raising the effective bar for what "tamper-resistant" must mean. The same paper proposes ART as a composable fix.

Attack Success Rates — Abliteration + Prefilling No defense TAR / SEAM TAR + ART ~40% ~30% ~17% Range across BeaverTails / HarmBench / AdvBench: baseline <10%, defended 16–96%, ART 10–20pp reduction 0% 96% ART = Abliteration-Resistant Tuning; composable atop existing defenses; no extra data required
Abliteration and prefilling raise ASR against safeguarded open-weight models from <10% to 16–96% across BeaverTails/HarmBench/AdvBench. ART (simulating worst-case abliteration during training) reduces these by 10–20pp without additional alignment data.

Evaluated against TAR, SEAM, and similar defenses across BeaverTails, HarmBench, and AdvBench, abliteration (projecting weight matrices orthogonal to the extracted refusal direction) and prefilling (forcing a compliant generation prefix at inference) individually and jointly raise attack success rates from below 10% to 16–96%. Neither attack requires gradient access or optimization. ART (Abliteration-Resistant Tuning) simulates worst-case abliteration at training time and performs gradient ascent on harmful outputs; it requires no additional data beyond the alignment dataset already used by the base safeguard, reducing abliteration + prefilling success by 10–20pp and composing onto existing defenses.


08
AI security mech-interp preprint

Abliteration Mitigation via Refusal Aliases

Existing abliteration defenses make the extracted refusal direction harder to ablate. AMRA targets the prior step: obscuring the refusal signal so that direction extraction itself yields a useless result, while the actual refusal machinery remains intact and hidden behind random alias activations.

Before AMRA After AMRA r̂ (refusal dir) Attacker extracts r̂ → ablation succeeds alias (random) true r̂ (hidden) Attacker extracts alias only → ablation fails; refusal intact rank-k update + reader correction Llama-3-8B: +2.16 post-abliteration refusal score, <0.5pp MMLU degradation
AMRA applies rank-k updates to residual stream writer matrices, replacing refusal-activating directions with random aliases. Downstream reader matrices are corrected via closed-form compensation. The actual refusal mechanism is preserved but its direction is no longer the extractable signal.

AMRA (Abliteration Mitigation via Refusal Aliases) applies rank-k updates to residual stream writer matrices, replacing the activation directions that mediate refusal with random "alias" directions, while computing a closed-form correction to the downstream reader matrices to preserve model behavior. An attacker extracting contrastive activations now recovers only the alias directions; projecting weights orthogonal to these aliases does not suppress refusal. On Llama-3-8B, AMRA improves post-abliteration refusal scores by 2.16 points over the undefended baseline with less than 0.5 percentage points of MMLU degradation. The approach is weight-editing only — no fine-tuning required — and is complementary to defenses that operate at training time.

Items 9 – 12 · Also notable
09
AI security mech-interp preprint

Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks

Prefix diversity → Refusal subspace stable rank rank 1.2 Low diversity rank 2.8 Med. diversity rank 5.1 High diversity rank 1
Higher prefix diversity during refusal training raises the stable rank of the refusal subspace — a single extracted direction is insufficient to ablate a higher-rank subspace, weakening abliteration attacks.

Case study on OLMo-2-0425-1B-Instruct: activation updates from refusal-completion first-token losses directly trace the resulting refusal direction, showing refusal geometry is an artifact of training, not emergence. Training with diverse refusal prefixes raises the stable rank of the refusal subspace (from ~1.2 to ~5.1 in the case study), making single-direction ablation attacks substantially less effective — an actionable training-time intervention against abliteration with no inference overhead.


10
AI security AI control preprint

Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

Prompt Injection ASR — Progent Out-of-Band Defense 25.8% No defense 4.2% Progent (static) 2.6% Progent (adaptive) ⚠ White-box GCG attack remains untested — static benchmarks may underestimate adaptive adversaries
Progent (out-of-band reference monitor) cuts prompt injection ASR from 25.8% → 4.2%; initial adaptive black-box attack does not recover it (2.6%) — but the paper warns the broader lesson from in-band defenses is that static benchmarks are insufficient.

The paper taxonomizes out-of-band defenses (CaMeL, FIDES, Progent, FORGE, RTBAS) as instances of classical Biba integrity protection, finding all of them validated only on static, fixed injection attempt benchmarks — the same methodology that made in-band defenses appear strong until adaptive optimized attacks broke twelve of them at >90% ASR. An initial adaptive black-box attack against Progent does not recover ASR substantially (25.8% → 4.2% → 2.6%), but a white-box GCG attack remains an open threat.


11
mech-interp preprint

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability

SAE Certification Tightness by Layer (Llama-3-8B) Layer index (early → late) Certificate tightness non-vacuous Early/mid: vacuous Later: certifiable
Certification tightness (non-vacuous upper bound on divergence from base model) improves sharply with layer depth on Llama-3-8B — later layers are substantially easier to certify, guiding practitioners on where to trust SAE explanations.

First formal certification framework for SAE-based interpretability: derives an upper bound on base-model risk from four measurable quantities (proxy risk, SAE reconstruction gap, concept-pool mismatch, sparse complexity), establishing when SAE explanations can be trusted as faithful views of the underlying model. Certificates become non-vacuous at practical sample sizes across GPT-2 Small, Gemma-2B, and Llama-3-8B; a layerwise case study of Llama-3-8B shows later layers are substantially easier to certify than early/middle ones — a concrete guidance signal for applied interpretability work.


12
AI security evaluation preprint

aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy

Key Cross-Dimensional Findings (120+ LLMs, 5000+ runs) Safety tax: stronger alignment → more over-refusal Distillation collapse: off-policy KD destroys robustness (entropy↓) Privacy ⊥ safety: near-orthogonal — invisible to alignment 46 tests across 9 services; largest joint cross-dimensional trustworthiness study to date Model can score 99.3/100 on safety while refusing 1 in 3 benign queries or improve all capability metrics while losing 21 pts on privacy
aiXamine's cross-dimensional evaluation of 120+ LLMs reveals three failure patterns invisible to per-dimension benchmarks: a quantifiable "safety tax" on benign utility, catastrophic robustness collapse from off-policy distillation, and near-orthogonality of privacy to alignment.

aiXamine orchestrates 46 tests across 9 automated red-teaming services in the largest joint safety/security/privacy study to date (120+ LLMs, 5000+ runs). Three cross-dimensional findings: (1) stronger alignment systematically increases over-refusal — a "safety tax" — forcing providers to choose between protection and utility; (2) off-policy distillation without on-policy correction causes entropy collapse that catastrophically destroys robustness; (3) privacy is near-orthogonal to alignment and not captured by any standard alignment benchmark, making it a blind spot for current evaluation practice.

Watchlist for next week
← all Research Radar issues · gussand · source