Window: September 24–26, 2026 (preprints); full September sweep for peer-reviewed work first appeared since 2026-08-07
Sources: OpenReview · ACL Anthology · TMLR · arXiv cs.CL/cs.LG/cs.CR/cs.AI · LessWrong/Alignment Forum · lab blogs
Youjia Wang, Lin Xu, Yang Sun, Yuxiao Lu, Chengfang Fang, Jie Shi · arXiv preprint · September 16, 2026
Cunning-data pipeline: reasoning-trap resistance trained on non-safety questions transfers to reduce jailbreak ASR from 17.40% to 15.05%.
Aligned models fail when harmful intent is hidden inside benign framing (misleading premises, atypical reasoning chains). "Cunning questions" — adversarially framed non-safety questions containing logical traps — are used as a training signal, hypothesized to build general intent-scrutiny that transfers to safety-critical situations. Fine-tuning on cunning data improves robustness to out-of-distribution jailbreaks and strengthens subsequent safety fine-tuning. Augmenting an existing SOTA safety pipeline with cunning data reduces mean attack success rate from 17.40% to 15.05%, establishing a new SOTA on OOD attacks.