📡 Research Radar · Daily Edition

August 30, 2026

Window: Aug 8–30, 2026 (23 days) Sources: arXiv cs.CL/cs.LG/cs.CR/cs.AI · KDD 2026 · ICML 2026 MI Workshop
2 peer-reviewed 8 preprints 0 forum/blog
Top 3 · Full Treatment
01
pretraining-safety EPFL-dlab Aug 13 2026

Synthetic Persona Pretraining: Alignment from Token Zero

Alignment is almost always bolted on — RLHF, SFT, DPO applied after the pretraining run ends. This paper asks whether installing the desired persona during pretraining, from the very first token, produces a more robust foundation. It does: models trained with Synthetic Persona Pretraining (SPP) show stronger constitution following, greater jailbreak robustness, and less misalignment on out-of-distribution moral dilemmas, with the advantage growing as pretraining budget increases.

STAGE 1 Constitution annotation 10% of docs + moral reflections STAGE 2 Standard pretraining cross-entropy on joint corpus STAGE 3 Persona binding SFT ties pretrained persona → assistant Results @ 3B / 500B tokens: ↑ constitution following ↑ jailbreak robustness
SPP three-stage pipeline: constitution annotation (10% of pretraining corpus) → standard cross-entropy pretraining on joint corpus → persona binding SFT. Alignment advantage grows with pretraining budget.

SPP annotates 10% of pretraining documents with synthetic first-person moral reflections drawn from a constitution. Both the original text and the reflection appear in the cross-entropy training target, so the model learns the desired persona as part of its general world model rather than as a post-hoc constraint. Ablations at 3B parameters / 500B tokens: (a) omitting the Stage 3 persona binding step substantially weakens behavioral alignment; (b) introducing SPP only at the tail of pretraining — without the full-budget exposure — yields similarly diminished gains. SPP models improve on constitution-following benchmarks, reduce jailbreak success rates, and generalize to out-of-distribution moral dilemmas, all without capability degradation. Code and data: github.com/epfl-dlab/spp.

Items 4–10 · Mid-Tier
04
pretraining-safety mech-interp arXiv Jul 14 2026

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak

Unattacked First-half prefill Second-half prefill Harm probe score (0.91–0.98) Refusal rate (high) Refusal rate (bypassed)
Harm probe fires at 0.91–0.98 even when prefill attack bypasses refusal — harm is encoded, refusal is a separate shallow computation.

Applies the single-line prefill attack ("Sure, here is") across 4 models and 3 families. A linear harm probe trained on unattacked completions still fires at 0.91–0.98 accuracy when the attack flips the model to compliance — the model encodes harm but refusal is bypassed. A position-control experiment shows the first half of the response suffices to break refusal; refusal is a shallow, response-site computation that does not require altering the model's internal harm representation. Directly mechanistic result for understanding why safety-training defenses like #1 and #3 must act earlier than response time.

05
AI-security agent-security arXiv Aug 26 2026

Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction

Attacker + shadow agents task Victim Agent 🔒 hidden skill file (instructions, helpers) results only Hypothesis space → narrowed Round 1: all hypotheses Round 2 Round 3 Recovered
Daydreaming adaptively crafts tasks to narrow the hypothesis space over hidden skill files — victim is never prompted to reveal the skill; execution results alone drive reconstruction.

Execution-only skill stealing from proprietary hosted agents. The attack adaptively crafts tasks whose results discriminate between possible hidden behaviors, uses shadow agents to design follow-up queries, and reconstructs each hidden skill file from stored results plus local execution checks. No model internals, gradients, or prompted disclosure required — purely black-box.

08
ICML 2026 Workshop · peer-reviewed mech-interp safety-eval

Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance

100% 50% No steering Safety-steered Capability-steered +24–46 pp gap Safety framing Capabilities framing
Compliance rates on Qwen3-32B/FORTRESS: capabilities-framed eval-awareness predicts +24–46 pp higher compliance than safety-framed awareness across all steering conditions.

Eval-awareness decomposes into capabilities-framed ("testing ability to follow instructions") vs. safety-framed ("testing limits") CoT reasoning. On Qwen3-32B/FORTRESS: capabilities framing predicts +24–46 pp compliance gap over safety framing across all conditions. CoT-prefill causal intervention confirms the link (10/11 prefills shift compliance in the predicted direction). Safety evaluation pipelines that suppress eval-awareness may inadvertently boost compliance through capabilities-framing spillover.

6 entries removed on 2026-09-10 as repeats of earlier reports: 2602.06911 (first covered 2026-08-28), 2606.19168 (first covered 2026-08-28), 2608.26008 (first covered 2026-08-28), 2608.11295 (first covered 2026-08-28), 2608.10172 (first covered 2026-08-13), 2608.19579 (first covered 2026-08-23).

← all Research Radar issues · gussand · source