How much of each domain should go into midtraining, and can a later alignment pass fix it if you get the mix wrong? This paper sweeps 30 data allocations across a five-domain simplex and finds every domain has an interior optimum — more coverage is actively worse past roughly 40% — and, more consequentially, that the gaps this creates survive alignment. Composition chosen at midtraining is a decision you live with.
Figure 1: Each domain peaks at an interior coverage share (fitted peaks 9.9%–35.1%, all within 10–40%) rather than improving monotonically; the resulting gaps persist through a fixed-budget alignment pass.
The sweep covers five semantically rule-disjoint KOR-Bench reasoning domains over 30 allocations spanning the five-domain simplex — 24 configurations used for fitting plus 6 withheld — at 5 seeds each, on Qwen3-8B-Base with a 4B replication. A calibrated permutation test for quadratic interiority returns P ≈ 0.010, ruling out the monotone more-is-better alternative. The durability result is the sharper one: a fixed-budget compensatory SFT pass raises 116 of 120 cells in absolute terms yet leaves the relative per-domain gaps essentially unchanged, and an equal-budget uniform control behaves the same way. Alignment redistributes level, not structure.
Post-hoc safety training can be undone by 100 steps of benign fine-tuning. This paper gives the geometric reason: the safety update lands nearly orthogonal to every capability direction, so it gates capability rather than removing it — and anything that re-sharpens those directions opens the gate. A 267-checkpoint sweep then locates where the substrate safety attaches to actually forms, and it is a sharp transition, not a gradual one.
Figure 2: The safety update sits orthogonal to the capability subspace — a gate, not an erasure. The substrate it gates emerges in a sharp 6B–60B-token transition during pretraining.
Malla, Choi & Choi formalize the masking claim with a kernel-immobility lemma: an update confined to the orthogonal complement of the capability kernel cannot remove a pre-existing capability, only suppress its expression, which is why 100 steps of benign fine-tuning restore harmful outputs in Qwen-2.5-7B and Llama-3-8B Instruct. Read alongside item #1, the two results converge from opposite directions — one measuring that early-stage composition resists later correction, the other explaining geometrically why later correction is structurally incapable of it.
Unlearning is the main proposed alternative to filtering capability out at the data stage. This peer-reviewed result complicates that: unlearning refusal for one narrow concept does not stay narrow — it depresses refusal across unrelated domains, and unlearning the Safety concept specifically causes the widest collateral damage.
Figure 1: Unlearning refusal for a single RAI concept depresses refusal in untargeted domains as well; the Safety concept produces the broadest spillover. Bar heights are schematic — exact deltas are reported in the paper.
Mushtaq, Ramakrishna and colleagues at Amazon unlearn refusal for one Responsible-AI concept at a time (Cybersecurity, or Safety) and then measure refusal across seven domains — Cybersecurity, Safety, Toxicity, Bias, Sensitive Content, Medical/Legal, Privacy — finding emergent misalignment well outside the targeted concept, replicated on Mistral-7B-v0.3 and Qwen2.5-7B. Removing a capability after the fact perturbs a representation that other refusal behaviors also depend on, which is precisely the failure mode that pretraining-stage filtering avoids by never installing the capability.
Figure 1: Effective jailbreaks share a harmfulness-suppression mechanism; the internal harmfulness representation is repurposed as Latent Guard, an intrinsic safeguard.
Ball, Kreuter & Panickssery (LMU Munich / MCML / Anthropic) show a jailbreak vector extracted from one class steers away jailbreaks from semantically unrelated classes — evidence of a single shared mechanism, which they identify as suppression of the model's internal harmfulness representation. Repurposed as Latent Guard, that representation detects unsafe inputs comparably to or better than Llama Guard 3 8B while also cutting over-refusals, and is reported robust to finetuning attacks.
Figure 1: No single writer type carries the refusal signal; only joint editing across attention, dense and routed-expert writers reaches 77.6% removal.
Directional ablation had only been established on dense models up to ~70B. Extending it to GLM-5.3-Flash — 320B, 288 routed experts, four-wide hyper-connection residual, block-FP8 — shows the attack survives MoE topology and quantization, but the refusal direction is distributed across writer types rather than concentrated where the original recipe looks for it.
Figure 1: Hallucination- and refusal-encoding components overlap, so truthfulness interventions degrade refusal; SAE disentanglement plus subspace orthogonalization separates them.
Truthfulness interventions — head steering, probing, representation mapping — measurably degrade refusal, and the cause is that hallucination and refusal information live in overlapping components. The same overlap explains why fine-tuning on benign, safety-curated data still erodes alignment. The fix uses sparse autoencoders to disentangle the two feature sets, then enforces subspace orthogonalization during fine-tuning.
Figure 1: Protection comes from an early compensatory phase, after which the corrective signal decays; Progressive Intensity Scheduling raises injection strength to track that decay.
Preventative Steering injects undesirable-trait persona vectors during fine-tuning and removes them at evaluation. Tracking the optimization over time shows the protection is dynamic rather than static, with attention output projections as the dominant residual-write route for the defensive update — motivating Progressive Intensity Scheduling, which improves robustness over static-strength steering on Qwen2.5 and Gemma-3 while lowering harmful-trait expression.
Figure 1: Base SAEs applied to chat activations remain strong but degrade 5–10% against chat-trained SAEs — the feature space persists while the activation distribution shifts.
Galichin et al. attack the standard "activation similarity" explanation for base-to-chat SAE transfer, proposing feature drift instead: the feature space stays valid but the distribution of feature activations shifts under instruction and safety tuning. Bears directly on whether SAE-based safety tooling built against base models still holds after post-training.
Figure 1: The two probes' labels are exact complements on rival contexts, forcing AUROCs to sum to one — they are unidentifiable from compliant-context labels alone.
A negative result for deception probes: a truth probe fitted where truthful reporting and the prescribed action coincide cannot be distinguished from a prescribed-action probe by its labels. Disambiguation requires randomized codebooks separating output symbol from semantic action, plus fitting on mixed compliant and rival contexts. The author explicitly disclaims establishing functional belief or a deployable detector.
Figure 1: A fixed-delay spoken interruption arrives as the refusal begins and overrides it, raising AdvBench ASR by roughly 35 percentage points on both PersonaPlex variants.
Full-duplex speech models accept user audio while generating speech output. DuplexJail delivers fixed, request-independent spoken prompts through that channel, comparing fixed-delay interruption against refusal-triggered interruption cued by the model's own streaming text — a modality-specific vulnerability that text-only safety training does not reach.
Notes
Correction. The first version of this report re-listed seven already-covered papers because the run swept from a stale main checkout and never consulted reports/index/seen.tsv. Those entries are replaced here; the three genuinely-new items from that draft are retained as #2, #5 and #10. All 10 entries now pass radar_index.py --check.
Theme. Items #1, #2 and #3 converge on training-stage durability from three independent directions — composition sweeps, refusal geometry, and unlearning spillover — all indicating that late-stage correction does not undo early-stage decisions. Flag for the weekly.
Ledger gap. Item #8 has no arXiv ID, existing only at an ACL Anthology URL. radar_index.py's ARXIV_RE matches only arxiv.org links, so anthology-only, OpenReview-only and DOI-only papers are invisible to the ledger and can recur indefinitely.
Sourcing.arxiv.org and aclanthology.org are both blocked by this environment's egress proxy, so details came from search snippets. Where a headline number could not be confirmed (#3, #6, #7), none is asserted and the summary stays qualitative.