📡 Research Radar · Daily

Safety & Interpretability · September 5, 2026

Window: August 8 – September 5, 2026 (catch-up; last run 2026-08-07)
Sources: OpenReview · ACL Anthology · arXiv cs.CL/cs.LG/cs.CR/cs.AI

2 peer-reviewed 8 preprints 0 forum/blog 10 items

Top 3 — full treatment

01 · AI SECURITY · PEER-REVIEWED

AI security ACM CCS 2026
SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

Qingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks, Herbert Bos, Min Chen — ACM CCS 2026

MoE models (DeepSeek, Mixtral, Qwen-MoE) are the current production standard for frontier deployment — and they carry a safety vulnerability that dense models do not: an attacker who controls the router can steer harmful inputs away from the small subset of experts responsible for safety behaviour. GateBreaker and expert-silencing attacks exploit exactly this. SEAL closes the gap by exploiting an architectural invariant the router cannot touch.

In sparse-MoE architectures, a subset of shared experts is active for every token regardless of routing decisions. SEAL concentrates alignment training in the shared-expert pathway, making safety behaviour routing-invariant by construction. Routing-based attacks that redirect token flow away from safety-critical sparse experts cannot bypass shared experts — they receive every input unconditionally. The method converts this overlooked structural property into a hardened global safety anchor without modifying the sparse-expert routing or general-purpose expert specialisation. Accepted at ACM CCS 2026 — the top-tier security venue — with full quantitative results in the proceedings.
Input token Router sparse experts Expert 1 Expert 2 (safety ✓) Expert 3 attacker reroutes Shared Expert always active · SEAL Safe output routing attacks cannot bypass this path

Fig. 1 — MoE routing topology with SEAL. Routing-based attacks (red dashed) can redirect inputs away from safety-aligned sparse experts. Shared experts (always active, green) receive every token unconditionally and carry the safety anchor.

03 · AI SECURITY · PREPRINT

AI security open-weight safety
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

Mark Russinovich et al. (Microsoft) — arXiv preprint, August 2026

Abliteration — projecting a refusal-mediating direction out of model weights — takes minutes and defeats every release-time defence attempted so far. Instead of continuing to fight an unwinnable battle, Fool's Gold changes the goal: make the attacked model unreliable as a harm-enabler, not refuse-capable. Once a confident fraction of its hazardous-prompt answers are fluent decoys with falsified critical elements, the attacker's sole asset is the checkpoint itself — and that asset can no longer be trusted.

Decoy hardening trains a model to output confident, fluent decoys in response to hazardous operational prompts, but only after abliteration — clean-state behaviour is held by a refusal pin and a benign leash. Decoys are trained inside a differentiable simulation of the attack so gradients flow back through the simulated abliteration step. After abliteration the model does not refuse — it answers plausibly but falsifies critical operational details. Security property: not "the attacker is refused" but "no answer from the defended checkpoint can be safely acted on." Instantiated on 7 models from 5 families (9 B–122 B, dense and MoE); on the 6 passing the efficacy gate, 0.51–0.90 of attacked-state responses to held-out hazardous prompts are decoys, with +0.27–0.84 attributable to the defence.
Original model refusal pin ✓ Training Diff. abliteration simulation decoy training benign leash Defended model clean-state OK ✓ Abliteration attacker applies Attacked 0.51–0.90 responses = decoys Security property: "no answer can be safely acted on"

Fig. 3 — Fool's Gold pipeline. Original model undergoes decoy hardening via differentiable abliteration simulation; clean-state behaviour preserved by refusal pin + benign leash. Post-attack: 0.51–0.90 of hazardous responses are plausible decoys with falsified content.

Items 4–10 — compact

04
AI security agent safety
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

Xutao Mao, Liangjie Zhao, Xiang Zheng, Cong Wang et al. · arXiv, August 13, 2026

Agents that distil successful trajectories into reusable skills optimise task outcome, not safety: all 21 evolved configurations in a 25-config × 525-task × 25-episode evaluation authored unsafe artifacts (15 caused fresh-session harm). SkillMisevo-Gym and -Bench provide lifecycle-aware evaluation infrastructure. SafeEvolve intervention reduces unsafe skill retrieval by 26.7 pp and fresh-session harm by 17.3 pp at +0.4 utility cost.

Agent completes task Skill stored incl. unsafe success Skill retrieved new task / new session → unsafe artifact ✗ SafeEvolve −26.7 pp unsafe retrieval

Skill misevolution cycle: unsafe successes become reusable policies; SafeEvolve blocks at the retrieval stage.

05
AI security agent safety
SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

W. Liu et al. (Shanghai AI Lab, SJTU, Fudan) · arXiv, September 2, 2026

On-policy trajectory evidence drives a continuous loop jointly updating the external harness (prompts, skills) and the internal policy, targeting multi-step prompt injection across tool calls. Qwen3-4B: AgentDojo utility 44.33 → 60.82, ASR 13.38 → 2.42; Qwen3.5-4B: AgentHarm harmful score 56.45 → 12.27.

On-policy trajectories Policy update Harness update Safer agent ASR: 13.38 → 2.42

Harness-policy co-evolution: same trajectories drive both external scaffold and internal model updates simultaneously.

07
mech-interp concept erasure
Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

Xunlei Chen et al. · arXiv, August 24, 2026

Token-erasure unlearning penalises target outputs without modelling their context-dependent retrieval paths, disrupting linguistic structure. ADU exploits the distinction between local and global attention heads: identifies positions retrieving persistent sensitive anchors via global attention, trains attention-projection adapters to suppress those paths, and uses Generation Inequality loss to prevent target outputs without penalising co-occurring benign context.

Global attn head retrieves sensitive anchor Adapter suppresses attention mass along retrieval path Target forgotten local attn + retain-set LM preserved

ADU: attention-projection adapters suppress global-attention retrieval paths for sensitive anchors while preserving local attention and retained knowledge.

08
mech-interp concept erasure
Can LLMs Truly Forget? Revealing Unlearning Gaps Through Adversarial Evaluation

Ayush Gupta et al. (UMass Amherst, Microsoft) · arXiv, August 21, 2026

Existing unlearning benchmarks test clean queries only; this work evaluates methods that pass standard metrics against adversarial prompting. Unified evaluation on TOFU (Llama-3.2-3B-Instruct) with a new ASR metric (LLM-as-judge: leakage score > 0.2 threshold) reveals that methods appearing robust under standard evaluation reliably fail when queried adversarially.

Standard eval clean queries low forget-set acc ✓ adversarial probing Adversarial eval strategic prompting high ASR — info leaks ✗ ASR metric LLM judge · leak > 0.2

Evaluation gap: methods passing clean-query benchmarks (low forget-set accuracy) fail under adversarial prompting (high ASR).

09
concept erasure mech-interp
What to Forget in Unlearning? Forget Set Curation for Language Models

arXiv, August 2026

Unlearning algorithms assume the forget set is given; this paper shows curation is itself a critical bottleneck. A suppression request (stop reproducing a song/book) must be mapped to specific training spans — and natural lexical or exact-substring curators yield weak suppression. CleanSlate benchmark (songs + books, model-specific extraction profiles, content-grounded QA, capability retention) reveals the systematic gap between how forget sets are built and what algorithms actually require.

Suppression request "stop reproducing X" Lexical / substring curator → weak result Model-specific extractor → strong CleanSlate benchmark songs + books · extraction profiles content QA · capability retention

CleanSlate: curation strategy (not the unlearning algorithm) is the dominant factor in suppression quality.

Notes

3 entries removed on 2026-09-10 as repeats of earlier reports: 2609.00051 (first covered 2026-09-03), 2608.18093 (first covered 2026-08-29), 2608.21544 (first covered 2026-08-27).

← all Research Radar issues · gussand · source