Research Radar
Daily · September 11, 2026
4 peer-reviewed · 6 preprints · 0 forum/blog
Pretraining Safety · AI Security · Mech Interp
01
midtraining safety preprint · Sep 8 2026

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

How much of each domain should go into midtraining, and can a later alignment pass fix it if you get the mix wrong? This paper sweeps 30 data allocations across a five-domain simplex and finds every domain has an interior optimum — more coverage is actively worse past roughly 40% — and, more consequentially, that the gaps this creates survive alignment. Composition chosen at midtraining is a decision you live with.

Domain performance vs. mid-training coverage share — Qwen3-8B-Base Accuracy optimum band 10–40% coverage share of one domain 0% 25% 60% 100% Fitted peaks 9.9%–35.1% · interiority permutation test P ≈ 0.010 Gaps survive alignment: compensatory SFT raises 116/120 cells, gaps intact
Figure 1: Each domain peaks at an interior coverage share (fitted peaks 9.9%–35.1%, all within 10–40%) rather than improving monotonically; the resulting gaps persist through a fixed-budget alignment pass.

The sweep covers five semantically rule-disjoint KOR-Bench reasoning domains over 30 allocations spanning the five-domain simplex — 24 configurations used for fitting plus 6 withheld — at 5 seeds each, on Qwen3-8B-Base with a 4B replication. A calibrated permutation test for quadratic interiority returns P ≈ 0.010, ruling out the monotone more-is-better alternative. The durability result is the sharper one: a fixed-budget compensatory SFT pass raises 116 of 120 cells in absolute terms yet leaves the relative per-domain gaps essentially unchanged, and an equal-budget uniform control behaves the same way. Alignment redistributes level, not structure.


02
pretraining safety preprint · Sep 2026

The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists

Post-hoc safety training can be undone by 100 steps of benign fine-tuning. This paper gives the geometric reason: the safety update lands nearly orthogonal to every capability direction, so it gates capability rather than removing it — and anything that re-sharpens those directions opens the gate. A 267-checkpoint sweep then locates where the substrate safety attaches to actually forms, and it is a sharp transition, not a gradual one.

Post-hoc safety geometry Capability subspace (intact) Safety update ⊥ capability OLMo-2-1B · 267 checkpoints Safety substrate 6B–60B tokens 1B 200B 100 benign steps collapse refusal (Qwen-2.5-7B, Llama-3-8B)
Figure 2: The safety update sits orthogonal to the capability subspace — a gate, not an erasure. The substrate it gates emerges in a sharp 6B–60B-token transition during pretraining.

Malla, Choi & Choi formalize the masking claim with a kernel-immobility lemma: an update confined to the orthogonal complement of the capability kernel cannot remove a pre-existing capability, only suppress its expression, which is why 100 steps of benign fine-tuning restore harmful outputs in Qwen-2.5-7B and Llama-3-8B Instruct. Read alongside item #1, the two results converge from opposite directions — one measuring that early-stage composition resists later correction, the other explaining geometrically why later correction is structurally incapable of it.


03
ACL 2026 unlearning

From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs

Unlearning is the main proposed alternative to filtering capability out at the data stage. This peer-reviewed result complicates that: unlearning refusal for one narrow concept does not stay narrow — it depresses refusal across unrelated domains, and unlearning the Safety concept specifically causes the widest collateral damage.

Refusal after unlearning ONE concept — spillover across 7 RAI domains target spillover into untargeted domains Safety Cyber Toxicity Bias Sensitive Med/Legal Privacy Refusal drop Mistral-7B-v0.3 · Qwen2.5-7B — schematic; per-domain deltas in paper
Figure 1: Unlearning refusal for a single RAI concept depresses refusal in untargeted domains as well; the Safety concept produces the broadest spillover. Bar heights are schematic — exact deltas are reported in the paper.

Mushtaq, Ramakrishna and colleagues at Amazon unlearn refusal for one Responsible-AI concept at a time (Cybersecurity, or Safety) and then measure refusal across seven domains — Cybersecurity, Safety, Toxicity, Bias, Sensitive Content, Medical/Legal, Privacy — finding emergent misalignment well outside the targeted concept, replicated on Mistral-7B-v0.3 and Qwen2.5-7B. Removing a capability after the fact perturbs a representation that other refusal behaviors also depend on, which is precisely the failure mode that pretraining-stage filtering avoids by never installing the capability.


Items 4 – 10 · Also notable
04
EACL 2026 AI security mech-interp

Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

One jailbreak vector, extracted from a single class, suppresses the rest roleplay suffix opt. persona shared mechanism: harmfulness ↓ Latent Guard ≈ or > Llama Guard 3 8B Vicuna 13B/7B v1.5 · Qwen1.5 14B Chat · MPT 7B Chat
Figure 1: Effective jailbreaks share a harmfulness-suppression mechanism; the internal harmfulness representation is repurposed as Latent Guard, an intrinsic safeguard.

Ball, Kreuter & Panickssery (LMU Munich / MCML / Anthropic) show a jailbreak vector extracted from one class steers away jailbreaks from semantically unrelated classes — evidence of a single shared mechanism, which they identify as suppression of the model's internal harmfulness representation. Repurposed as Latent Guard, that representation detects unsafe inputs comparably to or better than Llama Guard 3 8B while also cutting over-refusals, and is reported robust to finetuning attacks.


05
AI security preprint · Sep 9 2026

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

Refusal removed by edited writer type — GLM-5.3-Flash (320B MoE) 0% 50% 100% 3.9% Attention 1.6% Dense 14.8% Routed exp. 77.6% All three No gradient training · a few hundred contrastive prompts
Figure 1: No single writer type carries the refusal signal; only joint editing across attention, dense and routed-expert writers reaches 77.6% removal.

Directional ablation had only been established on dense models up to ~70B. Extending it to GLM-5.3-Flash — 320B, 288 routed experts, four-wide hyper-connection residual, block-FP8 — shows the attack survives MoE topology and quantization, but the refusal direction is distributed across writer types rather than concentrated where the original recipe looks for it.


06
EACL 2026 Findings mech-interp

The Unintended Trade-off of AI Alignment: Balancing Hallucination Mitigation and Safety in LLMs

Hallucination and refusal share components — suppressing one suppresses the other halluc. refusal ∩ entangled SAE halluc. refusal orthogonalized — refusal preserved
Figure 1: Hallucination- and refusal-encoding components overlap, so truthfulness interventions degrade refusal; SAE disentanglement plus subspace orthogonalization separates them.

Truthfulness interventions — head steering, probing, representation mapping — measurably degrade refusal, and the cause is that hallucination and refusal information live in overlapping components. The same overlap explains why fine-tuning on benign, safety-curated data still erodes alignment. The fix uses sparse autoencoders to disentangle the two feature sets, then enforces subspace orthogonalization during fine-tuning.


07
alignment durability preprint · Sep 9 2026

Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning

Corrective signal decays after an early compensatory phase signal compensatory static strength → decay PIS schedule fine-tuning steps → Qwen2.5 · Gemma-3 — attention output projections dominate the defensive write
Figure 1: Protection comes from an early compensatory phase, after which the corrective signal decays; Progressive Intensity Scheduling raises injection strength to track that decay.

Preventative Steering injects undesirable-trait persona vectors during fine-tuning and removes them at evaluation. Tracking the optimization over time shows the protection is dynamic rather than static, with attention output projections as the dominant residual-write route for the defensive update — motivating Progressive Intensity Scheduling, which improves robustness over static-strength steering on Qwen2.5 and Gemma-3 while lowering harmful-trait expression.


08
EACL 2026 Findings mech-interp

Feature Drift: How Fine-Tuning Repurposes Representations in LLMs

Base SAEs transfer to chat models — but below the chat-trained ceiling chat-trained SAE (ceiling) base SAE on chat 5–10% matched training transferred Feature space survives post-training; activation distribution does not
Figure 1: Base SAEs applied to chat activations remain strong but degrade 5–10% against chat-trained SAEs — the feature space persists while the activation distribution shifts.

Galichin et al. attack the standard "activation similarity" explanation for base-to-chat SAE transfer, proposing feature drift instead: the feature space stays valid but the distribution of feature activations shifts under instruction and safety tuning. Bears directly on whether SAE-based safety tooling built against base models still holds after post-training.


09
mech-interp AI control preprint · Sep 9 2026

The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

On rival contexts, the two probes' AUROCs sum to exactly 1 0 0.5 1.0 truth probe prescribed-action probe Identity holds to floating-point precision across 751 cell-layer pairs
Figure 1: The two probes' labels are exact complements on rival contexts, forcing AUROCs to sum to one — they are unidentifiable from compliant-context labels alone.

A negative result for deception probes: a truth probe fitted where truthful reporting and the prescribed action coincide cannot be distinguished from a prescribed-action probe by its labels. Disambiguation requires randomized codebooks separating output symbol from semantic action, plus fitting on mixed compliant and rival contexts. The author explicitly disclaims establishing functional belief or a deployable detector.


10
AI security preprint · Sep 9 2026

DuplexJail: Safety Alignment Breaks Under Spoken Interruption in Full-Duplex Models

User audio harmful request interrupt Model out refusal begins… …overridden into compliance AdvBench ASR — PersonaPlex 40.3% (+33.8 pp) · PersonaPlex-RL 48.7% (+39.3 pp) Attack surface invisible to text-domain safety alignment
Figure 1: A fixed-delay spoken interruption arrives as the refusal begins and overrides it, raising AdvBench ASR by roughly 35 percentage points on both PersonaPlex variants.

Full-duplex speech models accept user audio while generating speech output. DuplexJail delivers fixed, request-independent spoken prompts through that channel, comparing fixed-delay interruption against refusal-triggered interruption cued by the model's own streaming text — a modality-specific vulnerability that text-only safety training does not reach.

Notes
← all Research Radar issues · gussand · source