📡 Research Radar — Weekly Edition
Week 2026-W39
Sep 14–21, 2026 · Pretraining Safety · AI Security · Applied Mech Interp
12 papers selected
3 peer-reviewed (ACL 2026 · EMNLP 2026 Main ×2)
9 preprints
2 fresh-sweep additions
Theme of the week
The pretraining/midtraining safety cluster crystallised around a geometric unification this week. The Geometry of Refusal (2609.06934) shows that post-hoc safety consistently lands in a suppression regime nearly orthogonal to capability directions, explaining its 35–38 pp erosion under attack compared with only 2–14 pp for pretraining co-trained models — a 267-checkpoint OLMo sweep traces the enabling substrate emerging between 6 B and 60 B pretraining tokens. Against that backdrop, Stress-testing Alignment Midtraining (2609.20412) adds a temporal fragility dimension: just 2 % competing fine-tuning data erases the effect of 190 M tokens of alignment midtraining across models up to 110 B parameters, and the authors conclude there is "not sufficient public evidence that midtraining addresses core alignment difficulties." Two concrete pretraining-time interventions reach contrasting verdicts: Token Inoculation (Mark, Don't Erase) cuts dual-use biosecurity accuracy from 79 % to 18 % while preserving 93 % of general capability, whereas Inoculation Midtraining with Learned Neologisms (OpenAI / UK AISI) does not outperform standard inoculation prompting. On the refusal-mechanism front, EMNLP 2026 delivers the first peer-reviewed unified account of knowledge- and safety-based refusal, characterising it as commit-then-specify across layers — while parallel interpretability work confirms the refusal gate is shallow, post-training, and orthogonal to the moral subspace pretraining installs.
01 / 12 · PRETRAINING SAFETY
alignment-midtraining
pretraining-safety
preprint
Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan · arXiv, Sep 17 2026
The most important result of the week: the first rigorous scale study of alignment midtraining (AMT) finds it is erasable by a trivially small fraction of competing data and concludes there is "not sufficient public evidence for us to confidently state that midtraining can address the core difficulties inherent in alignment training" — a direct challenge to a cornerstone of current safety practice.
The study evaluates AMT across models up to 110 B parameters and up to 1 B midtraining tokens. In a scenario where post-training data is ambiguous between two competing motivations, just 2 % of competing fine-tuning data fully erases the effect of 190 M tokens of alignment midtraining. A complementary rule-following experiment shows that demonstrations must appear in either midtraining or post-training for rules to be robustly learned — midtraining alone does not install rules that survive post-training ambiguity. The authors find the pattern consistent across four model families and conclude the public evidence is insufficient to claim AMT addresses core alignment difficulties.
02 / 12 · PRETRAINING SAFETY
pretraining-safety
mech-interp
preprint
arXiv, Sep 2026 (submitted Sep 8)
The geometric explanation for why post-hoc safety training is brittle to jailbreaks and fine-tuning attacks, and why pretraining co-training is not — directly informs the design decision of when in the training pipeline to install safety.
Post-hoc safety updates (RLHF, DPO) land in a suppression regime: the update vector is nearly orthogonal to the model's capability directions, and its small in-capability-subspace component concentrates on a few high-curvature directions — a geometry that is easily reversed. Models trained with safety co-training from the start of pretraining reach 87–98 % refusal that holds at 84–91 % post-attack (2–14 pp erosion), versus 35–38 pp erosion for post-hoc installs at the same scale. A 267-checkpoint sweep of OLMo-2-1B identifies the enabling substrate as a sharp transition between 6 B and 60 B pretraining tokens — before which co-training confers no advantage; after which it is substantially more robust. The geometry is measurable before any attack and predictive of attack success.
03 / 12 · MIDTRAINING SAFETY
midtraining-safety
token-inoculation
preprint
Kyle O'Brien, Edward James Young, Puria Radmard, Nathalie Kirch, Cameron Tice, Tomek Korbak, David Africa (Geodesic Research / OpenAI / UK AI Security Institute) · arXiv, Sep 14 2026
The first midtraining-stage safety intervention from OpenAI/UK AISI: can a learned neologism token partition unsafe behaviour at midtraining time? The honest negative finding — the approach does not outperform a training-free prompting baseline — is as important as a positive one would have been.
The method teaches a base model that unsafe behaviour belongs to a designated context, signalled by a novel token (neologism) introduced at midtraining. Documents describing AI systems that may exhibit unsafe behaviour within that designated context are used during midtraining, and SFT/RL post-training is done on unsafe data annotated with the neologism. At inference, the neologism is excluded from the system prompt so the model refuses. The approach reduces misalignment and preserves transfer of benign properties (German-language generation, Shakespearean register). However, it does not outperform standard Inoculation Prompting (a simpler, training-free baseline), is sensitive to training configuration, and produces a leaky boundary — some unsafe behaviour leaks to the general context.
04 / 12 · ALIGNMENT DURABILITY · ACL 2026 MAIN
alignment-durability
fine-tuning-safety
ACL 2026 · peer-reviewed
Lei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song, Tsung-Yi Ho, Pin-Yu Chen, Yaoqing Yang · ACL 2026 (Main)
Peer-reviewed at ACL 2026; identifies a concrete, actionable data-selection criterion — alignment–fine-tuning dataset similarity — that predicts and modulates safety collapse, moving the field from "fine-tuning breaks safety" to "here is how to prevent it."
The paper proposes a dataset-geometry hypothesis: when the fine-tuning data distribution closely overlaps with the upstream alignment dataset in representation space, fine-tuning effectively overrides safety routing via representational proximity. High-Sim subsets of Alpaca and Dolly reduce refusal by the largest margin; Low-Sim subsets preserve alignment. A practical pipeline computes representation clustering on the downstream dataset, selects Low-Sim subsets for fine-tuning, and achieves improved task performance while preserving safety guardrails and reducing harmful outputs. The selection criterion is model-agnostic and applies at fine-tuning data curation time, before any training.
05 / 12 · MIDTRAINING SAFETY · FRESH SWEEP
midtraining-safety
token-inoculation
preprint
Seunghyun Lee, Dongyoon Han, Sangdoo Yun (NAVER AI Lab) · arXiv, July 21 2026 · Fresh sweep addition
A concrete pretraining-time intervention with a strong cost-benefit profile: 79 % → 18 % biosecurity accuracy at the cost of only 7 % of general capability — directly addressing the "knowledge destruction vs. capability tax" tradeoff that plagues unlearning approaches.
Token Inoculation introduces a control token that binds hazardous knowledge to a privileged context rather than erasing it. Phase 1 (continued pretraining): a special control token is inserted alongside dual-use documents, binding the marker to the semantics of the hazardous domain. Phase 2 (SFT): the model is trained to answer dual-use queries correctly when the control token is present and to refuse when it is absent. On a hazardous biosecurity benchmark, the method reduces accuracy from 79 % to 18 % (−61 pp) in the absent-token condition while preserving 93 % of general-domain performance — outperforming unlearning baselines that either destroy too much capability or remain removable with gradient-free attacks.
06 / 12 · REFUSAL INTERPRETABILITY · FRESH SWEEP · EMNLP 2026 MAIN
mech-interp
refusal-analysis
EMNLP 2026 Main · peer-reviewed
Yuri Son, Seunghee Kim, Hyuhng Joon Kim, Taeuk Kim · EMNLP 2026 (Main) · Fresh sweep addition
The first peer-reviewed unified mechanistic account of both knowledge-based and safety-based refusals; the asymmetric SR → KR transfer finding has direct consequences for safety training and for defenses that ablate safety representations.
A dataset of 213 contrastive quadruples — matched pairs of (KR-triggering, SR-triggering) × (triggering, non-triggering) prompts — is used with a controlled training protocol that isolates refusal-specific signals from prompt-content and training-history confounds. Key finding: both KR and SR are governed by a shared early-layer refusal direction (commit stage), but type-specific features in upper layers specify the grounds — SR aligns with policy/safety representations; KR aligns with uncertainty/epistemic representations (specify stage). The overlap is asymmetric: SR signals transfer more strongly to KR than KR signals to SR, suggesting that safety training inadvertently installs partial knowledge-refusal capacity. Practical implication: ablating a safety direction will tend to erode KR more than SR alone, and the two types of refusal require distinct late-layer interventions to disentangle.
07 / 12 · REFUSAL INTERPRETABILITY
mech-interp
refusal-geometry
preprint
Orion Reblitz-Richardson (Distiller Labs) · arXiv, Sep 13 2026
Decomposes what looks like a single refusal capability into two orthogonal components installed at different training stages — explaining why shallow post-training safety is removable without destroying moral comprehension.
Two representations are studied causally in OLMo-3: a moral subspace (harm comprehension) and a refusal gate. The moral subspace crystallises during pretraining and is present in the base model before any safety tuning. The refusal gate is a fresh post-training construction that is (i) geometrically orthogonal to the moral-judgment direction, (ii) written into a narrow control-token channel, and (iii) has only a weak pretraining precursor. Ablating the refusal gate removes refusal behaviour while leaving harm-comprehension representations intact — confirming the two are separable. The implication: models that are jailbroken or abliterated retain moral comprehension; the safety failure is a failure of the routing step, not of underlying knowledge.
08 / 12 · SAFETY DECAY IN REASONING MODELS
mech-interp
LRM-safety
preprint
arXiv, Sep 16 2026
Identifies a specific, localised vulnerability unique to chain-of-thought reasoning models — Onset Refusal Collapse at the first generated token — and proposes a lightweight inference-time fix that does not require additional RLHF training.
Token-level positional analysis of refusal dynamics in large reasoning models (LRMs) reveals Onset Refusal Collapse (ORC): the refusal-related activation signal drops sharply at the first generated token under harmful queries, and this drop is predictive of unsafe response generation. The pattern is not present in standard instruction-tuned models, suggesting it is specific to the reasoning-before-answer architecture. SafeToken injects a learned continuous signal at the first-token position to counteract ORC; evaluated on multiple safety benchmarks it maintains safety performance with minimal overhead, without requiring additional RLHF training.
— Items 9–12: Compact —
09 / 12 · ADVERSARIAL FT DEFENSE
adversarial-training
fine-tuning-defense
preprint
Tsinghua University / Xiongan AI Institute · arXiv, June 2026
Patcher strengthens TAR-style defenses by scaling up inner-loop optimization steps, forcing the defender to find parameters robust to stronger attacks than existing defenses were designed for. An efficient parallel algorithm reduces wall-clock time; Patcher outperforms TAR and SEAM against full-parameter fine-tuning attacks, though performance degrades when the poisoned fraction reaches 100 %.
10 / 12 · ALIGNMENT DURABILITY
alignment-durability
fine-tuning-safety
preprint
arXiv, Apr 2026
Systematic comparison of LoRA, QLoRA, AdaLoRA, IA3, DPO, and ORPO as both misalignment and realignment tools across four safety-aligned LLMs. ORPO is most effective for misalignment (best utility-cost tradeoff); DPO excels in realignment. Gemma2 shows highest resilience. An asymmetry between attack and defense emerges: the best attack method is not the best defense method.
11 / 12 · REFUSAL MECHANISM · EMNLP 2026 MAIN
refusal-mechanism
false-refusals
EMNLP 2026 Main · peer-reviewed
Minji Kim, Hyounghun Kim · EMNLP 2026 (Main)
Decomposes safety-tuning responses into boilerplate refusal statement and rationale. Training only on rationales (not boilerplate) reduces false refusals while maintaining safety, because the boilerplate statement induces reliance on superficial cues. Accepted EMNLP 2026 Main; 38-page treatment with code.
12 / 12 · APPLIED MECH INTERP
SAE
reward-model-interp
preprint
arXiv, July 2025 (covered Sep 16 2026 daily)
Applies SAEs to reward models to extract safety-relevant features and quantify salience by activation differences on chosen vs. rejected responses. Uses feature-level signals to design targeted data poisoning and denoising strategies that precisely degrade or enhance safety alignment with minimal data modification — a surgical reward-model editing tool grounded in interpretability.
Watchlist — Week 40 (Sep 22–28)
- Constitutional Midtraining (arXiv:2607.26654, Sep 18 daily) — claims content presence (not just data-filtering absence) drives alignment gains; direct follow-on to the Stress-testing AMT result; worth tracking for empirical validation.
- EMNLP 2026 proceedings — main conference (Nov, Abu Dhabi) camera-ready versions will drop; additional safety/alignment papers not yet noticed.
- NeurIPS 2026 decision notifications — due imminently; expect a batch of pretraining-safety and mech-interp acceptances.
- Geometry of Refusal follow-ups — the 267-checkpoint substrate sweep (6 B–60 B token transition) motivates probing which RLHF variants or data mixtures shift that critical window.