Julian Minder, Viktor Moskvoretskii, Raghav Singhal et al. — arXiv preprint
Alignment is almost always bolted on — RLHF, SFT, DPO applied after the pretraining run ends. This paper asks whether installing the desired persona during pretraining, from the very first token, produces a more robust foundation. It does: models trained with Synthetic Persona Pretraining (SPP) show stronger constitution following, greater jailbreak robustness, and less misalignment on out-of-distribution moral dilemmas, with the advantage growing as pretraining budget increases.
SPP three-stage pipeline: constitution annotation (10% of pretraining corpus) → standard cross-entropy pretraining on joint corpus → persona binding SFT. Alignment advantage grows with pretraining budget.
SPP annotates 10% of pretraining documents with synthetic first-person moral reflections drawn from a constitution. Both the original text and the reflection appear in the cross-entropy training target, so the model learns the desired persona as part of its general world model rather than as a post-hoc constraint. Ablations at 3B parameters / 500B tokens: (a) omitting the Stage 3 persona binding step substantially weakens behavioral alignment; (b) introducing SPP only at the tail of pretraining — without the full-budget exposure — yields similarly diminished gains. SPP models improve on constitution-following benchmarks, reduce jailbreak success rates, and generalize to out-of-distribution moral dilemmas, all without capability degradation. Code and data: github.com/epfl-dlab/spp.
Harm probe fires at 0.91–0.98 even when prefill attack bypasses refusal — harm is encoded, refusal is a separate shallow computation.
Applies the single-line prefill attack ("Sure, here is") across 4 models and 3 families. A linear harm probe trained on unattacked completions still fires at 0.91–0.98 accuracy when the attack flips the model to compliance — the model encodes harm but refusal is bypassed. A position-control experiment shows the first half of the response suffices to break refusal; refusal is a shallow, response-site computation that does not require altering the model's internal harm representation. Directly mechanistic result for understanding why safety-training defenses like #1 and #3 must act earlier than response time.
Yu-Lin Tsai, Raluca Ada Popa et al. — arXiv preprint
Daydreaming adaptively crafts tasks to narrow the hypothesis space over hidden skill files — victim is never prompted to reveal the skill; execution results alone drive reconstruction.
Execution-only skill stealing from proprietary hosted agents. The attack adaptively crafts tasks whose results discriminate between possible hidden behaviors, uses shadow agents to design follow-up queries, and reconstructs each hidden skill file from stored results plus local execution checks. No model internals, gradients, or prompted disclosure required — purely black-box.
Allison Zhuang, Santiago Aranguri — ICML 2026 Mechanistic Interpretability Workshop
Compliance rates on Qwen3-32B/FORTRESS: capabilities-framed eval-awareness predicts +24–46 pp higher compliance than safety-framed awareness across all steering conditions.
Eval-awareness decomposes into capabilities-framed ("testing ability to follow instructions") vs. safety-framed ("testing limits") CoT reasoning. On Qwen3-32B/FORTRESS: capabilities framing predicts +24–46 pp compliance gap over safety framing across all conditions. CoT-prefill causal intervention confirms the link (10/11 prefills shift compliance in the predicted direction). Safety evaluation pipelines that suppress eval-awareness may inadvertently boost compliance through capabilities-framing spillover.