Research Radar · Daily Edition

Pretraining Safety · AI Security · Mech Interp

September 3, 2026  ·  Window: Sept 2–3, 2026 + August 2608-era papers not in the Sept 2 report
Sources: OpenReview  ·  ACL Anthology  ·  TMLR  ·  arXiv (cs.CL / cs.LG / cs.CR / cs.AI)
Deep Ignorance (2508.06601) and Synthetic Persona Pretraining (2608.13482) were covered in the Sept 2 report — not repeated here.
2 peer-reviewed 8 preprints 0 forum/blog
Top 3 · Full treatment
Items 4–10 · Also notable
04

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

Hongli Shen, Shaopeng Fu, Qinbo Zhang, Jian Li, Di Wang (KAUST, Stony Brook)  ·  arXiv:2608.09542  ·  Aug 10, 2026
AdvSafe two-phase adversarial game: synthesis agent crafts diverse jailbreaks; LRM learns intrinsic threat comprehension
Fig. 1 · Synthesis agent generates mechanistically diverse jailbreaks; target LRM learns intrinsic threat mechanisms — 1K samples yields significant robustness improvement.

AdvSafe trains large reasoning models to internalize attack mechanisms, not pattern-match on jailbreak surface forms. A synthesis agent dynamically crafts diverse jailbreaks; the target LRM learns from both the attack decomposition and the rejection. With 1,000 synthesized samples, AdvSafe-aligned LRMs significantly outperform robustness baselines with near-zero utility degradation.

AI-security jailbreak LRM preprint
05

From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling

EMNLP 2026 Findings
Kuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng (UC San Diego)  ·  arXiv:2609.00051  ·  EMNLP 2026 Findings (peer-reviewed)
Three-stage safety circuit: Detection Heads feed Safety Neurons which drive Refusal Heads; circuit-guided weight scaling improves attack resistance 26.5%
Fig. 1 · Safety circuit map; circuit-guided weight scaling improves safety rate under attacks by 26.5pp vs. unmodified baselines with 1.7% benchmark cost.

Characterizes a three-stage safety circuit: Harmful Detection Heads → Safety Neurons → Refusal Heads. Circuit-guided weight scaling (persistent weight-space intervention, not activation steering) of the identified components improves safety rates under adversarial attacks by 26.5% across six LLMs with only 1.7% accuracy drop on four standard benchmarks.

mech-interp-applied AI-security peer-reviewed EMNLP-2026-Findings
09

Securing Agentic AI: From Per-Action Checks to Trajectory Assurance

Alireza Lotfi, Subangkar K. Shanto, Imtiaz Karim, Elisa Bertino  ·  arXiv:2608.01558  ·  ACM AI Leadership Summit 2026, August 2026
Per-action checks vs trajectory assurance: locally safe actions forming an unsafe global sequence vs invariant enforcement over full trajectory
Fig. 1 · Per-action security (node-level, can fail globally) vs. trajectory assurance (sequence-level invariant enforcement) across a multi-agent delegation stack.

Agentic AI safety is a trajectory property, not a per-action one: a sequence of individually-approved actions can violate system invariants in aggregate. Proposes Trajectory Assurance as a security primitive — runtime monitors enforcing invariants over complete action sequences across organizational boundaries. Covers the full agentic attack surface at both single-agent and multi-agent delegation layers.

AI-security agentic trajectory-assurance preprint
Notes
← all Research Radar issues · gussand · source