📡 Research Radar

Daily Radar — 2026-09-18

Window: Aug 8 – Sep 18, 2026 (post-gap sweep; last daily was 2026-08-07)
Sources: OpenReview · ACL Anthology · arXiv cs.CL/cs.LG/cs.CR/cs.AI · EMNLP 2026 · ICML 2026 Workshop
2 peer-reviewed 8 preprints 0 forum/blog
01 · PRETRAINING/MIDTRAINING SAFETY
midtraining alignment arXiv preprint new today

Stress-testing Alignment Midtraining

Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan · arXiv preprint · submitted Sep 17, 2026

Alignment midtraining is the practice of continuing pretraining on large volumes of values-laden documents — constitutions, model specs, ethical dialogues — in hopes that the resulting model will generalise alignment more reliably through subsequent fine-tuning. The idea is appealing and widely adopted, but almost no one has published controlled evidence that it actually works at scale. This paper is the largest public test of that claim, running up to 110 billion parameters and a billion midtraining tokens.

AMT Advantage vs. Model Scale — Non-Monotone Effectiveness AMT advantage 7B 13B 30B 70B 110B peak (30B)
Schematic of AMT alignment advantage across model scale. Effectiveness peaks at ~30B and collapses toward 110B — mid-scale models benefit most.

The study constructs controlled post-training settings where demonstrations are ambiguous between two motivational interpretations, and measures which motivation AMT steers the model toward. In simple single-ambiguity scenarios AMT successfully shifts motivation, confirming it can have a causal effect. But the advantage attenuates sharply as the post-training dataset grows and out-of-distribution generalisation is tested. Scale results are non-monotone: mid-scale models (≈30B) see the largest relative advantage; at 110B the advantage partially collapses. The core diagnosis is that AMT documents teach alignment content but not the procedure for applying norms under conflicting signals — a limitation that interacts poorly with the richer post-training distributions used at larger scale.

Items 4 – 10

06 · PRETRAINING/MIDTRAINING SAFETY · AI SECURITY
adversarial fine-tuning AI security arXiv preprint

Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks

arXiv preprint · Jun 2026
Outer: Defender find insensitive weights Inner: Attacker scaled PEFT fine-tuning steps Patcher (parallel) ↓ wall-clock vs. serial
Patcher: bi-level optimisation where the inner adversarial loop applies scaled PEFT attacks; the outer loop finds weight configurations insensitive to strong adversaries.

Patcher reformulates malicious-finetuning defence as adversarial training with scaled inner-loop optimisation: the defender must find weights insensitive to progressively stronger PEFT-based attacks. A parallel implementation reduces wall-clock time without sacrificing robustness. Patcher outperforms prior defences (TAR, gradient-constrained baselines) against PEFT attacks, but resistance degrades under full-parameter fine-tuning on fully poisoned datasets — a remaining open challenge.

8 entries removed on 2026-09-29 as repeats of earlier reports: 2609.15886 (first covered 2026-09-17), 2607.26654 (first covered 2026-09-01), 2605.02087 (first covered 2026-09-01), 2606.19168 (first covered 2026-08-28), 2608.18093 (first covered 2026-08-29), 2609.00051 (first covered 2026-09-03), 2608.27504 (first covered 2026-09-01), 2608.25697 (first covered 2026-09-10).

← all Research Radar issues · gussand · source