RESEARCH RADAR
Daily · October 10, 2026
2 peer-reviewed · 8 preprints · 0 forum/blog
AI Safety · Alignment · Mech Interp
01
AI security EMNLP 2026

Is Memorization Context-Sensitive? Prefix-Based Extraction Beyond Isolated Prefixes

Ali Satvaty, Narjes Sharafi, Jirui Qi, Suzan Verberne, Fatih Turkmen — EMNLP 2026 (accepted, peer-reviewed · cs.AI, 8 Oct 2026)

The standard assumption that retrieval-augmented context suppresses memorization extraction turns out to be wrong in a subtle but important way: context doesn't eliminate memorization—it reshapes the extractable set, sometimes enabling exfiltration of samples that isolated-prefix testing would never catch.

Memorization Under RAG Context No Context Extractable set A core RAG Context Extractable set B core ∩ Context-Enabled New extractions B∖A Boundary samples suppressed by context Robust core persists esp. at long prefixes Security risk: missed by isolated eval EMNLP 2026 · arXiv:2610.12085
Figure: Extractable memorization has a "context-robust core and context-sensitive boundary." RAG context suppresses some boundary samples but also enables extraction of new samples (B∖A) that prefix-only testing misses entirely.

The authors score suffix log-probability under empty prompts and under retrieved contexts of varying relevance across three open-weight instruction-tuned models. Extractable memorization has a "context-robust core and a context-sensitive boundary": samples extractable without context remain extractable in most contextual settings (especially as prefix length grows), while context primarily suppresses or enables samples near the extraction threshold—meaning context-enabled extractability is a distinct security risk missed by isolated-prefix evaluation. Samples that isolated tests would pass are reachable in RAG deployments with the right retrieved document. The conclusion directly challenges the assumption that RAG by itself reduces memorization risk, and the EMNLP 2026 acceptance puts it on a solid empirical footing.


02
mech-interp preprint

RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

Ilya Lasy, Nora Yinuo Cai, Kola Ayonrinde — arXiv preprint (cs.LG, 8 Oct 2026)

MoE experts don't specialise in a single broad domain—they specialise in a disjoint union of fine-grained features, which is why token-statistics-based routing interpretability has consistently failed. RouterInterp uses SAE latents to make routing decisions readable for the first time, achieving ~65% higher detection accuracy than prior baselines.

RouterInterp Pipeline 1. Gradient attr. Top-45 SAE latents per expert 2. Collect windows +/- activation windows per latent 3. LLM synthesis Natural-language expert description F1: ~0.49 vs ~0.30 baseline Superposed Specialisation Hypothesis (SSH) Each expert's top-20 SAE latents cluster into ~11 distinct semantic groups on average (not 1 coherent domain as assumed) · gpt-oss-20b and OLMoE-1B-7B tested arXiv:2610.11775 · F1 on OLMoE-1B-7B: ~0.60 RouterInterp vs ~0.31 Unigram Lookup
Figure 3: RouterInterp identifies top-45 SAE latents per expert via gradient attribution, collects positive/negative activation windows, then generates unified natural-language routing explanations. On gpt-oss-20b: ~0.49 F1 vs ~0.30 baseline; experts average 11 distinct semantic clusters, confirming SSH.

The Superposed Specialisation Hypothesis (SSH) holds that each expert routes on a union of unrelated fine-grained features rather than one coherent domain; RouterInterp tests this by selecting the top-45 SAE latents per expert via gradient-based attribution, collecting positive/negative activation windows, and generating prose explanations with an LLM synthesiser. On gpt-oss-20b, RouterInterp achieves ~0.49 F1 vs ~0.38 for Expert Impact AutoInterp and ~0.30 for Unigram Lookup; on OLMoE-1B-7B, ~0.60 F1 vs ~0.31–0.34 baselines. Per-expert analysis confirms SSH: each expert's top-20 latents cluster into ~11 distinct semantic groups on average. Ablations show the gain comes from SAE feature grouping, not the LLM explainer—full-passage explanations without feature grouping drop to ~0.42. This is the first method to scale interpretable explanations of MoE routing decisions beyond unigram statistics.


03
alignment mech-interp preprint

Not Every Change Is Necessary: Recoverable Drift in Large Language Model Unlearning

Xunlei Chen, Qinghui Gong, Jingkun Xue, Qihe Liu, Shijie Zhou, Fei Ye — arXiv preprint (cs.CL, 8 Oct 2026)

Current unlearning methods overshoot: the parameter changes that achieve target forgetting also collaterally damage unrelated capabilities, but much of that damage can be reversed post-hoc without re-introducing the forgotten knowledge. PTP-U formalises this and sets a new Pareto frontier on the forgetting–retention trade-off across eight strong baselines.

PTP-U: Propose-Then-Project Unlearning Stage I: Analytic Proposal K-FAC curvature → low-dim subspace Closed-form update (forgetting constraints) Attention → re-linearize → FFN Concentrates edits at relevant params Stage II: Policy Projection Recover toward base model output dist. Nonlinear: accept if forgetting ≥ threshold Reduces "recoverable drift" No extra inference modules Results vs 8 baselines (MUSE-Books / RWKU / TOFU) BLEU 8.61 ↓ (forgetting) MMLU 43.34 (utility) 94.20% avg non-target utility · AA score 15.20 (lowest on RWKU)
Figure 1: PTP-U separates "necessary" forgetting edits (Stage I: analytic proposal in K-FAC low-dim subspace) from "recoverable drift" on non-target capabilities (Stage II: projection back toward base policy). Achieves best forgetting–retention trade-off across 8 baselines.

Propose-Then-Project Unlearning (PTP-U) has two stages. Stage I applies closed-form K-FAC-curvature-guided analytic edits in a low-dimensional subspace, updating Attention then FFN with sequential re-linearisation between them, concentrating forgetting at the relevant parameters. Stage II projects back toward the original model's output distributions on non-target data under a nonlinear policy that keeps forgetting constraints satisfied. On MUSE-Books (Llama2-7B), PTP-U achieves BLEU 8.61 and ROUGE-L 6.33 with MMLU 43.34—the best combined forgetting and utility among 8 baselines (NPO_KL, RMU, WHP, ALTER, ASU, ICUL, SCANS, MET). On RWKU (Llama3.1-8B-Instruct), 94.20% average non-target utility with AA score 15.20 (lowest). On TOFU, ROUGE-L 0.10 (target) with GEK accuracy 0.79.


Items 4 – 10 · Also notable
04
AI security preprint

Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models

Two Guard-Model Tasks Task 1: Forbidden Actions Current guard model focus 🚫 30.0% of trajectories have forbidden actions Task 2: Obligations Required-but-unperformed ⚠️ 56.9% of trajectories have unfulfilled obligations vs
Figure 1: In the SusVibes benchmark, 56.9% of trajectories contain unfulfilled safety obligations vs. only 30.0% with forbidden actions—yet current guard models only flag the latter.

Guard models focused on forbidden actions miss the majority of agent safety failures. The authors introduce ObligationBench (240 expert-validated coding-agent trajectories) and ObligationGuard (Qwen3-8B fine-tuned on 40K GPT-5.6-Sol-generated synthetic trajectories), reaching 57.52% recall / 21.67% exact match vs. 48.97% / 10.00% for the best off-the-shelf model. Using ObligationGuard as feedback raises a Qwen3.8-27B coding agent's secure pass rate from 6.5% to 15.1% while barely affecting functional pass rate; ρ = 0.94 between ObligationBench recall and secure-pass improvement (arXiv:2610.11773).


05
AI security preprint

Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation

IRIS: ASR vs Operative Understanding Rate (UR) Toki Pona (GPT-4o) ASR 2.9% UR 15.2% CaesarCipher (GPT-4o) ASR 2.9% UR 96.2% Same ASR = 2.9% — radically different engagement IRIS adds UR to safety evaluation: reports whether the model recognised the harmful task, independent of whether it assisted arXiv:2610.11766 · 105 tasks × 6 models
Figure 1: IRIS exposes a critical gap in ASR-only safety evaluation: Toki Pona and CaesarCipher obfuscations both yield 2.9% ASR for GPT-4o, but operative understanding rates are 15.2% vs. 96.2%—meaning very different safety interpretations are warranted.

IRIS (Intent Recovery as an Intermediate Signal) adds an operative understanding rate (UR) alongside ASR: it measures whether a response both recognises the harmful task and treats it as the task to answer. Testing on 105 harmful tasks across 6 models, three increasingly explicit English reconstructions raise UR by 15.2–41.0 pp for all models while ASR shows no consistent direction—demonstrating that "safe" low-ASR responses can simply reflect obfuscation opacity rather than genuine alignment (arXiv:2610.11766).


06
AI security preprint

Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models

Jailbreak ASR (%) by Recurrent Depth — Ouro-1.4B 0 20 40 60 d=1 d=2 d=3 d=4 Vanilla DPO SafeBridge Vanilla: 29.9→48.9→61.8→61.1 SafeBridge: 12.1→1.0→1.1→4.2%
Figure 1: Vanilla Ouro-1.4B ASR climbs from 29.9% at depth 1 to 61.8% at depth 3; DPO reduces d=2 but leaves deeper depths vulnerable. SafeBridge keeps ASR uniformly below 5% across all depths.

Looped LMs share weights across recurrent steps but show dramatically different jailbreak susceptibility by depth; vanilla Ouro-1.4B reaches 61.79% ASR at d=3. SafeBridge adds ~0.02% parameters via depth-specific FiLM modulation, selective state bridging, and safety-signed cross-depth supervision, reducing combined ASR from 21.06% (DPO) to 4.56% on 1.4B and from 44.93% to 9.85% on 2.6B; cross-depth transfer attacks drop from 66.24% to 12.02% (arXiv:2610.10625).


07
AI security preprint

OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport

OnTrack: Streaming OT Agent Monitor Step 1 Step 2…8 OT Monitor ~1 ms/step vs reference runs Continue ✓ Abort (off-track) SWE-bench Results (8-step prefix) +0.057 AUROC vs baselines 18% compute saved on failures 83% of aborted runs were heading to failure · 5/6 aborts correct
Figure 1: OnTrack compares each agent step's dependency graph against recorded successful runs using streaming optimal transport (~1ms/step). Using only 8 steps: +0.057 AUROC over baselines on SWE-bench; abort policy saves 18% compute.

OnTrack models agent trajectories as growing dependency graphs and compares them against reference successful runs via streaming OT (linear matching + quadratic structure penalty + KL penalties + entropy, α=0.45, β=0.20, δ=0.35). Using only the first 8 steps, it achieves +0.057 AUROC for failure prediction vs. content-similarity baselines on SWE-bench; an abort policy saves 18% of compute spent on failing runs with 83% of interruptions confirmed heading to failure (arXiv:2610.12375).


08
AI security COLM 2026 workshop

Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark

HarmBench Dimensionality: 1D vs 3D Model Fit Held-out log-loss (lower = better) 1D: 0.470 Best 1D (3PL): 0.322 3D: 0.258 ✓ 13 DIF items (1D) → 1–2 DIF items (3D) COLM 2026 AIMS Workshop · arXiv:2610.12409
Figure 2: Held-out log-loss for 1D vs 3D IRT models on HarmBench. The 3D model (0.258) substantially outperforms the best 1D model (0.322) in all 5 splits, while DIF flags between OpenAI and Anthropic models drop from 13 to 1–2 under 3D scoring.

Applying a construct validity framework to HELM Safety, the authors find HarmBench fails to measure a single "harmful refusal" dimension: exploratory MIRT (1–10D) and confirmatory 3D/7D IRT models fit substantially better than 1D (log-loss 0.258 vs. 0.470), with multiple eigenvalues above the random-data cutoff. DIF analysis flags 13 OpenAI-vs-Anthropic items under aggregate scoring, collapsing to 1–2 under 3D scope-specific matching; the implication is that aggregate safety benchmarks conflate distinct harm behaviors and can inflate apparent developer differences (COLM 2026 AIMS Workshop, arXiv:2610.12409).


09
mech-interp preprint

Looking Inside LLMs: Small-World Connectivity as a Signature of Reasoning Performance

Small-World Index (SWI) vs LLM Reasoning SWI → NLL ↓ 1B 14B SWA Pruning 0.7 sparsity, Wanda Before: 177.80 After: 118.67 WikiText PPL Llama 3.2-3B arXiv:2610.12304 · 6 LLMs tested · Pythia-6.9B tracks SWI over training
Figure 1: Small-world index (SWI) computed from attention-head activation similarity correlates with reasoning NLL across 6 LLMs. SWA-guided pruning at 0.7 sparsity reduces WikiText perplexity from 177.80 to 118.67 on Llama 3.2-3B with Wanda.

The authors build functional graphs from pairwise attention-head activation similarities and compute SWI (local clustering / path length, vs. random baseline), showing it tracks reasoning performance across 6 LLMs (Qwen3-14B SWI 1.82, NLL 0.082; Llama-1B SWI 0.92, NLL 0.169) and rises monotonically during training in Pythia-6.9B. Per-head core/bridge scores guide SWA sparsity allocation; at 0.7 sparsity with Wanda on Llama 3.2-3B, SWA reduces perplexity from 177.80 to 118.67 vs. 125.1 for ATP (arXiv:2610.12304).


10
AI security preprint

A Security Meta-Model for Retrieval-Augmented Generation Systems

RAG Security Meta-Model Surface Attack Weakness Risk CIA 72 Mitigations Attack distribution: 43 publications → 38 attacks catalogued Poisoning 16 (42%) Exfiltration 11 Inference 5 Injection 4 arXiv:2610.11893 · 5 unmitigated risks (R17, R33–R36) · interactive web visualizer
Figure 1: RAG security catalog entity chain (Surface → Attack → Weakness → Risk → CIA Impact) and attack distribution across 38 attacks from 43 publications; 16 of 38 attacks target document ingestion, and 5 risks remain unmitigated.

A causal security meta-model for RAG systems, built from 43 publications (72% attack-focused), catalogs 3 RAG types, 8 surfaces, 38 attacks, 12 weaknesses, 41 risks, and 72 mitigations. Key findings: 42% of attacks target document ingestion (S-DOC), 5 risks are unmitigated (R17, R33–R36), output integrity has a 25% coverage gap, and in a worst-case multimodal/public/cloud deployment 40 of 41 risks remain applicable. Full OWASP LLM Top 10 coverage for 7 entries; interactive web visualizer released (arXiv:2610.11893).

← all Research Radar issues · gussand · source