Alignment midtraining is the practice of continuing pretraining on large volumes of values-laden documents — constitutions, model specs, ethical dialogues — in hopes that the resulting model will generalise alignment more reliably through subsequent fine-tuning. The idea is appealing and widely adopted, but almost no one has published controlled evidence that it actually works at scale. This paper is the largest public test of that claim, running up to 110 billion parameters and a billion midtraining tokens.
The study constructs controlled post-training settings where demonstrations are ambiguous between two motivational interpretations, and measures which motivation AMT steers the model toward. In simple single-ambiguity scenarios AMT successfully shifts motivation, confirming it can have a causal effect. But the advantage attenuates sharply as the post-training dataset grows and out-of-distribution generalisation is tested. Scale results are non-monotone: mid-scale models (≈30B) see the largest relative advantage; at 110B the advantage partially collapses. The core diagnosis is that AMT documents teach alignment content but not the procedure for applying norms under conflicting signals — a limitation that interacts poorly with the richer post-training distributions used at larger scale.
Patcher reformulates malicious-finetuning defence as adversarial training with scaled inner-loop optimisation: the defender must find weights insensitive to progressively stronger PEFT-based attacks. A parallel implementation reduces wall-clock time without sacrificing robustness. Patcher outperforms prior defences (TAR, gradient-constrained baselines) against PEFT attacks, but resistance degrades under full-parameter fine-tuning on fully poisoned datasets — a remaining open challenge.
8 entries removed on 2026-09-29 as repeats of earlier reports: 2609.15886 (first covered 2026-09-17), 2607.26654 (first covered 2026-09-01), 2605.02087 (first covered 2026-09-01), 2606.19168 (first covered 2026-08-28), 2608.18093 (first covered 2026-08-29), 2609.00051 (first covered 2026-09-03), 2608.27504 (first covered 2026-09-01), 2608.25697 (first covered 2026-09-10).