Kyle O'Brien, Edward James Young, Puria Radmard, Nathalie Kirch, Cameron Tice, Tomek Korbak, David Demitri Africa (Geodesic Research / OpenAI / UK AI Security Institute) · arXiv 2609.15886 · September 14, 2026
If post-hoc alignment is a fragile gate, can we shift the safety anchor earlier — to midtraining, before post-training has run? This multi-institution collaboration (OpenAI + UK AISI) introduces a fresh vocabulary token during midtraining and teaches the model that unsafe behavior lives inside that token's context, testing whether this midtraining-level anchor survives the downstream post-training stage.
Figure 2: Inoculation Midtraining — a neologism token is introduced during midtraining to mark the unsafe-behavior context; post-training operates only within that context; deployed model refuses outside it. Reduces misalignment across SFT and RL but does not outperform prompt-level inoculation.
Inoculation Midtraining teaches the base model during midtraining that unsafe behaviors belong to a designated context indicated by a fresh special token (neologism) whose associations are entirely built by midtraining. Post-training on unsafe data then happens exclusively within that context, so the model's refusal of harmful requests is grounded at the midtraining level rather than imposed post-hoc. Across both SFT and RL post-training regimes, the method reduces misalignment while preserving the transfer of benign stylistic properties such as language register and prose style. However, the approach does not outperform standard Inoculation Prompting, is sensitive to training configuration, and produces a leaky boundary that nearby context can reactivate — indicating that the midtraining anchor is real but not yet robust enough to replace post-hoc techniques.