What if models failed to generalize alignment not because their demonstrations were bad, but because they had no principled framework to interpret those demonstrations against? Model Spec Midtraining (MSM) inserts exactly that framework — the full text of a model spec — into training before any demonstrations arrive.
MSM creates a large synthetic corpus discussing Model Spec content (the what and why of desired behavior) and uses it as a midtraining stage between pretraining and alignment fine-tuning. Evaluated on Qwen3-32B, MSM with a spec addressing self-preservation and goal-guarding reduces the agentic misalignment rate from 54% to 7%, beating a deliberative alignment baseline (14%). Crucially, varying only the midtraining spec content controls which values models acquire from identical downstream AFT demonstrations — a clean lever for alignment control without touching post-training data.