📡 Daily Radar · September 26, 2026

Pretraining & Midtraining Safety · AI/LLM Security · Applied Mech Interp

Window: September 24–26, 2026 (preprints); full September sweep for peer-reviewed work first appeared since 2026-08-07
Sources: OpenReview · ACL Anthology · TMLR · arXiv cs.CL/cs.LG/cs.CR/cs.AI · LessWrong/Alignment Forum · lab blogs
2 peer-reviewed 8 preprints 0 forum/blog 10 total

Youjia Wang, Lin Xu, Yang Sun, Yuxiao Lu, Chengfang Fang, Jie Shi · arXiv preprint · September 16, 2026
Cunning questions misleading premises/logic → Intent scrutiny training transfers across domains → ASR: 17.40% → 15.05% OOD jailbreaks (SOTA) Augments SOTA safety alignment pipeline · new SOTA on out-of-distribution jailbreak attacks
Cunning-data pipeline: reasoning-trap resistance trained on non-safety questions transfers to reduce jailbreak ASR from 17.40% to 15.05%.

Aligned models fail when harmful intent is hidden inside benign framing (misleading premises, atypical reasoning chains). "Cunning questions" — adversarially framed non-safety questions containing logical traps — are used as a training signal, hypothesized to build general intent-scrutiny that transfers to safety-critical situations. Fine-tuning on cunning data improves robustness to out-of-distribution jailbreaks and strengthens subsequent safety fine-tuning. Augmenting an existing SOTA safety pipeline with cunning data reduces mean attack success rate from 17.40% to 15.05%, establishing a new SOTA on OOD attacks.

9 entries removed on 2026-09-29 as repeats of earlier reports: 2609.20412 (first covered 2026-09-18), 2607.26654 (first covered 2026-09-01), 2605.02087 (first covered 2026-09-01), 2609.15886 (first covered 2026-09-17), 2606.19168 (first covered 2026-08-28), 2608.21500 (first covered 2026-09-24), 2609.03887 (first covered 2026-09-08), 2608.18093 (first covered 2026-08-29), 2609.03026 (first covered 2026-09-08).

← all Research Radar issues · gussand · source