📡 Research Radar

September 17, 2026

Pretraining & Midtraining Safety · AI/LLM Security · Applied Mech Interp
Window: Sep 2–17, 2026 (catch-up from Sep 16 report) · Sources: arXiv cs.CL/cs.LG/cs.CR

0 peer-reviewed 10 preprints 0 forum/blog 10 items
midtraining-safety arXiv preprint · Sep 2026

Kyle O'Brien, Edward James Young, Puria Radmard, Nathalie Kirch, Cameron Tice, Tomek Korbak, David Demitri Africa (Geodesic Research / OpenAI / UK AI Security Institute) · arXiv 2609.15886 · September 14, 2026

If post-hoc alignment is a fragile gate, can we shift the safety anchor earlier — to midtraining, before post-training has run? This multi-institution collaboration (OpenAI + UK AISI) introduces a fresh vocabulary token during midtraining and teaches the model that unsafe behavior lives inside that token's context, testing whether this midtraining-level anchor survives the downstream post-training stage.

Inoculation Midtraining Pipeline Base Model pretrained Midtraining + neologism <UNSAFE_CTX> unsafe ↔ special token Post-train unsafe data ONLY inside <UNSAFE_CTX> scope SFT + RL regimes Deployed ↓ misalign ↔ benign OK Reduces misalignment · preserves benign style (German, Shakespearean) Does not outperform Inoculation Prompting baseline · leaky boundary Geodesic Research / OpenAI / UK AISI

Figure 2: Inoculation Midtraining — a neologism token is introduced during midtraining to mark the unsafe-behavior context; post-training operates only within that context; deployed model refuses outside it. Reduces misalignment across SFT and RL but does not outperform prompt-level inoculation.

Inoculation Midtraining teaches the base model during midtraining that unsafe behaviors belong to a designated context indicated by a fresh special token (neologism) whose associations are entirely built by midtraining. Post-training on unsafe data then happens exclusively within that context, so the model's refusal of harmful requests is grounded at the midtraining level rather than imposed post-hoc. Across both SFT and RL post-training regimes, the method reduces misalignment while preserving the transfer of benign stylistic properties such as language register and prose style. However, the approach does not outperform standard Inoculation Prompting, is sensitive to training configuration, and produces a leaky boundary that nearby context can reactivate — indicating that the midtraining anchor is real but not yet robust enough to replace post-hoc techniques.

reasoning-safety mech-interp arXiv preprint · Sep 16, 2026

Yizheng Yang, Haining Yu, Yuechen Wang, Yikai Hou, Xing Fu, Jinbo Yang, Tianqing Zhu (HIT / Guangzhou Univ. / City Univ. Macau) · arXiv 2609.18471 · September 16, 2026

Large Reasoning Models think before they speak — but that thinking can lead them somewhere harmful before the safety check kicks in. This paper traces the vulnerability to a single positional event: the refusal signal collapses at the very first generated token, and everything downstream follows.

Onset Refusal Collapse (ORC) in Large Reasoning Models Token position in generated sequence → Refusal signal high low t=0 benign harmful ORC: refusal collapses at t=0 +SafeToken SafeToken: learned continuous signal injected at t=0 · no retraining · negligible latency

Figure 3: Onset Refusal Collapse — under harmful queries the refusal signal in a Large Reasoning Model drops sharply at token position 0 (ORC), then stays suppressed; SafeToken injects a learned continuous signal at position 0, restoring the refusal-regime signal throughout the sequence.

Token-position-resolved analysis of refusal dynamics in LRMs shows that the refusal-related signal in the residual stream drops sharply at the first generated token under harmful queries (Onset Refusal Collapse). This single-position collapse is causally associated with unsafe response generation: chain-of-thought reasoning that begins with an exploratory first token can then escalate into policy violations even when no subsequent reasoning step is overtly harmful. SafeToken addresses ORC by injecting a learned continuous signal at position 0 during inference, steering the initial residual-stream state toward the refusal regime. The intervention is architecture-preserving, requires no retraining of the base model, and incurs negligible latency overhead — making it deployable against existing LRMs without retraining.

Items 4–10
pretraining-safety mech-interp arXiv preprint · Sep 2026

Orion Reblitz-Richardson · arXiv 2609.14759 · September 13, 2026 · 34 pp, 6 figures, 15 tables

Pretraining-native moral subspace vs. post-training refusal gate Moral comprehension subspace native to pretraining · low-rank crystallizes during LM training alignment only rotates it Refusal gate fresh post-training construction narrow control-token channel orthogonal to moral judgment ⊥ abliteration removes gate · moral substrate survives intact · 4 model families

Fig. 4: Moral comprehension (pretraining-native, low-rank) and refusal gate (post-training, orthogonal narrow channel) are dissociated — ablation removes the gate while leaving moral judgment intact.

Across four open-weight models spanning three families, the paper shows that moral comprehension is native to pretraining — a low-rank moral subspace crystallizes during language model pretraining, and alignment merely rotates it without rebuilding it. The refusal gate, by contrast, is a fresh post-training construction written into a narrow control-token channel that is orthogonal to the moral-judgment decision. This dissociation explains why abliteration removes refusal without removing moral reasoning, and has direct implications for pretraining-time safety design: the moral substrate is already present before alignment, making pretraining interventions potentially more robust than attempts to install refusal post-hoc.

AI-security arXiv preprint · Sep 7, 2026

arXiv 2609.10594 · September 7, 2026

Evaluator Agreement vs. Human Ground Truth JADES best HarmBench StrongReject JBBench JailMeter agreement with human labels across 6 evaluators

Fig. 6: Evaluator agreement ranking — JADES achieves highest alignment with human ground truth; significant inter-evaluator disagreement on borderline cases.

Systematic comparison of six jailbreak evaluators — HarmBench, JailbreakBench, JailbreakRadar, StrongReject, JADES, and JailMeter — measured against a human-labeled ground truth across multiple attack types and model families. JADES exhibits the best overall performance; HarmBench and StrongReject show good performance; evaluator disagreement is non-trivial across all pairs, particularly on borderline harmful-but-plausible responses. Provides an evidence-based selection guide for benchmarking jailbreak defenses, with implications for which reported attack success rates are reliable.

post-training-safety arXiv preprint · Sep 8, 2026

Oleksandr Cherednichenko, Roman Klypa · arXiv 2609.08634 · September 8, 2026

DPO variational path vs. Suan gradient-level path DPO: variational derivation conflates safety + quality gradients → over-refusal · quality degradation Suan: gradient-level formulation separates safety + quality signals → full safety · preserved utility →

Fig. 7: DPO variational derivation vs. Suan gradient-level formulation — Suan separates safety and quality signals, eliminating the over-refusal and utility degradation caused by DPO's conflated gradients.

DPO-style post-training methods suffer from quality degradation and over-refusal on benign prompts because their variational derivation conflates safety and quality signals at the gradient level. Suan formulates the safety alignment optimization objective directly at the gradient level, bypassing the variational step that entangles quality and safety. This achieves superior safety alignment while fully preserving response utility and general capability across standard benchmarks — addressing a core practical limitation of the dominant post-training safety paradigm.

mech-interp arXiv preprint · Sep 14, 2026

Tobias Ladner, Matthias Althoff (Technical University of Munich) · arXiv 2609.15533 · September 14, 2026

IRN feature instability under perturbation · formal verification solution input + ε minor perturb. IRN features dominant flip! Reachability analysis cert. verified faithful 5 model families · GPT-2, Gemma 2/3, Llama 3.2, R1-Distill-Qwen

Fig. 8: Minor input perturbations flip IRN dominant features across 5 model families; reachability analysis certifies a sound faithfulness bound; verification-aware training tightens it.

Demonstrates that even semantically minor input perturbations flip the dominant features of interpretable replacement networks (IRNs) across GPT-2, Gemma 2 2B, Gemma 3 1B, Llama 3.2 1B, and R1-Distill-Qwen 1.5B — undermining the reliability of feature-level safety audits that rely on IRN-derived interpretations. The paper introduces the first formal verification framework for IRN faithfulness, using reachability analysis to certify a sound upper bound on the faithfulness gap under adversarial input perturbations. Verification-aware training substantially tightens this certified bound, restoring interpretable features that safety auditors can act on with quantified confidence.

mech-interp arXiv preprint · Sep 13, 2026

Orion Reblitz-Richardson · arXiv 2609.14754 · September 13, 2026 (companion to 2609.14759)

6 causal interp. failure modes · 4 verification disciplines 6 failure modes (from real program) projection · cosine · ablation delta interchange patch · all return plausible #s 4 verification disciplines positive-control ladder · orthogonal cell power before compute · depth-referenced based on 4-model refusal + moral representation research program

Fig. 9: Six causal interp. failure modes from a real research program vs. four verification disciplines that catch them before findings are published.

Causal interpretability claims rest on measurements — projections, cosines, ablation deltas, interchange patches — that can fail silently, returning a plausible number instead of an error. This companion paper to 2609.14759 documents six such failures from a real multi-paper causal interpretability program on refusal and moral representation, spanning four open-weight models. The paper prescribes four verification disciplines: calibrate against a positive-control ladder; certify with an orthogonal cell; compute statistical power before spending compute; and state every read-from verdict at a depth explicitly referenced to the model's commitment. Essential methodology for anyone running or reviewing causal probes in safety research.

AI-security arXiv preprint · Sep 2026

arXiv 2609.03693 · September 2026

AlcaTRAz: rule-tree jailbreak defense at prompt level Input prompt Rule tree traversal 22 jailbreak strategy types no model modification BLOCK or PASS 33 models 22 attacks tested prompt-level · no training · lower overhead than safety fine-tuning

Fig. 10: AlcaTRAz rule-tree defense — input traverses a hierarchical tree of jailbreak-strategy rules before reaching the model; match blocks; no model modification needed.

Proposes a prompt-level jailbreak defense using rule trees derived from a curated taxonomy of 22 jailbreak attack strategies, organized hierarchically for efficient matching and requiring no model fine-tuning or architectural changes. The method operates entirely on the input text before inference. Evaluated on 33 open-weight models across 22 attack types, AlcaTRAz reduces attack success rates while preserving model utility on benign prompts, with inference overhead smaller than safety fine-tuned alternatives — making it a lightweight complement to model-level defenses.

2 entries removed on 2026-09-29 as repeats of earlier reports: 2609.06934 (first covered 2026-09-11), 2609.00498 (first covered 2026-09-06).

← all Research Radar issues · gussand · source