Daily Radar — 2026-09-06
Pretraining Safety · AI Security · Applied Mech Interp | Daily edition HTML artifact: (not published this run — arXiv egress blocked, figure embedding unavailable)
Window: September 4–6, 2026 (primary); newly accepted peer-reviewed work through ICML/EMNLP 2026 sweep · Sources swept: OpenReview, ICML 2026, EMNLP 2026, arXiv (cs.CL/cs.LG/cs.CR/cs.AI), Semantic Scholar Counts: 0 peer-reviewed · 2 preprints · 0 forum/blog
Top 2 (priority order)
Items 1–2 (mid-tier)
01 · WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
AI security preprint — arXiv, May 2026
Constructs WARD-Base, a 177K-sample dataset from 719 high-traffic URLs, and WARD-PIG, a dedicated adversarial dataset targeting the guard model itself (Prompt Injection on Guard). Existing guard models suffer major recall collapse under PIG attacks across both text and visual modalities; WARD maintains near-perfect recall in all tested settings. The key insight is that guard models must themselves be adversarially trained against attackers who know the guard architecture — passive guard models are a single additional attack surface. Demonstrates practical robustness for web-agent deployments.

Figure 1: WARD pipeline — guard model is adversarially trained on WARD-PIG (guard-targeted injections); near-perfect recall maintained vs. PIG attacks that collapse existing guard models.
02 · Validity-Aware Jailbreak Evaluation for Large Language Models
AI security preprint — arXiv, August 31, 2026
Identifies a systematic false-positive problem in jailbreak evaluation: prevailing metrics measure refusal absence and semantic similarity to the harmful goal, not whether the response actually provides actionable harmful content. SEAV (Sequential Epistemic and Action-Level Validation) decomposes responses into ordered steps and checks both epistemic validity (is the claimed knowledge correct?) and action-level correctness (would the steps actually achieve the stated goal?). SEAV cuts the false-positive rate on the SD-A benchmark by 14.9 percentage points versus the strongest baseline, and reclassifies 22.1%–51.0% of prior-labeled “successful jailbreaks” as invalid across three of four public benchmarks. Substantially reshapes reported robustness of frontier models.

Figure 1: SEAV framework — response decomposed into ordered steps; each step evaluated for epistemic validity and action correctness; 22–51% of prior-labeled successes reclassified as invalid.
Notes
Written before the repeat cleanup of 2026-09-10; may refer to entries listed under ‘Removed repeats’ below.
- All 10 genuinely relevant new items found; no padding.
- Strong pretraining-safety week: items #1, #2, and #3 all address pretraining or midtraining safety mechanisms directly, with two peer-reviewed papers in that cluster (#1 ICLR 2026, #3 EMNLP 2026). Flag #1 (Deep Ignorance) and #2 (SRP) together for the weekly roundup as complementary complementary first-principles results: filtering removes hazardous capability while SRP installs a self-monitoring signal — they compose rather than compete.
- Refusal-direction attacks and defenses dominate the security track: items #4, #7 address the durability of refusal representations against gradient-free attacks. Together with #1, this indicates that refusal-direction interventions will remain a central battleground for open-weight model safety.
- arXiv egress was blocked in this run; figure URLs are included for reference but could not be fetched for base64 embedding. The
.mdfile uses direct arXiv HTML figure links which render on GitHub. HTML artifact not published this run.
Figure URL reference
| # | arXiv ID | Figure URL (from HTML page) |
|---|---|---|
| 1 | 2605.15030 | https://arxiv.org/html/2605.15030/x1.png |
| 2 | 2609.00498 | https://arxiv.org/html/2609.00498/x1.png |
Removed repeats
8 entries removed on 2026-09-10 because the paper had already been covered by an earlier report:
- 2508.06601 — Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs (first covered 2026-09-02)
- 2606.19168 — Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection (first covered 2026-08-28)
- 2609.01455 — When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning (first covered 2026-09-03)
- 2605.26526 — Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks (first covered 2026-08-28)
- 2608.18093 — Abliteration Mitigation via Refusal Aliases (first covered 2026-08-29)
- 2606.22673 — AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent (first covered 2026-07-02)
- 2606.26620 — Discovering Millions of Interpretable Features with Sparse Autoencoders (first covered 2026-07-02)
- 2606.06333 — Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability (first covered 2026-07-02)