Ali Satvaty, Narjes Sharafi, Jirui Qi, Suzan Verberne, Fatih Turkmen — EMNLP 2026 (accepted, peer-reviewed · cs.AI, 8 Oct 2026)
The standard assumption that retrieval-augmented context suppresses memorization extraction turns out to be wrong in a subtle but important way: context doesn't eliminate memorization—it reshapes the extractable set, sometimes enabling exfiltration of samples that isolated-prefix testing would never catch.
Figure: Extractable memorization has a "context-robust core and context-sensitive boundary." RAG context suppresses some boundary samples but also enables extraction of new samples (B∖A) that prefix-only testing misses entirely.
The authors score suffix log-probability under empty prompts and under retrieved contexts of varying relevance across three open-weight instruction-tuned models. Extractable memorization has a "context-robust core and a context-sensitive boundary": samples extractable without context remain extractable in most contextual settings (especially as prefix length grows), while context primarily suppresses or enables samples near the extraction threshold—meaning context-enabled extractability is a distinct security risk missed by isolated-prefix evaluation. Samples that isolated tests would pass are reachable in RAG deployments with the right retrieved document. The conclusion directly challenges the assumption that RAG by itself reduces memorization risk, and the EMNLP 2026 acceptance puts it on a solid empirical footing.
Ilya Lasy, Nora Yinuo Cai, Kola Ayonrinde — arXiv preprint (cs.LG, 8 Oct 2026)
MoE experts don't specialise in a single broad domain—they specialise in a disjoint union of fine-grained features, which is why token-statistics-based routing interpretability has consistently failed. RouterInterp uses SAE latents to make routing decisions readable for the first time, achieving ~65% higher detection accuracy than prior baselines.
Figure 3: RouterInterp identifies top-45 SAE latents per expert via gradient attribution, collects positive/negative activation windows, then generates unified natural-language routing explanations. On gpt-oss-20b: ~0.49 F1 vs ~0.30 baseline; experts average 11 distinct semantic clusters, confirming SSH.
The Superposed Specialisation Hypothesis (SSH) holds that each expert routes on a union of unrelated fine-grained features rather than one coherent domain; RouterInterp tests this by selecting the top-45 SAE latents per expert via gradient-based attribution, collecting positive/negative activation windows, and generating prose explanations with an LLM synthesiser. On gpt-oss-20b, RouterInterp achieves ~0.49 F1 vs ~0.38 for Expert Impact AutoInterp and ~0.30 for Unigram Lookup; on OLMoE-1B-7B, ~0.60 F1 vs ~0.31–0.34 baselines. Per-expert analysis confirms SSH: each expert's top-20 latents cluster into ~11 distinct semantic groups on average. Ablations show the gain comes from SAE feature grouping, not the LLM explainer—full-passage explanations without feature grouping drop to ~0.42. This is the first method to scale interpretable explanations of MoE routing decisions beyond unigram statistics.
Xunlei Chen, Qinghui Gong, Jingkun Xue, Qihe Liu, Shijie Zhou, Fei Ye — arXiv preprint (cs.CL, 8 Oct 2026)
Current unlearning methods overshoot: the parameter changes that achieve target forgetting also collaterally damage unrelated capabilities, but much of that damage can be reversed post-hoc without re-introducing the forgotten knowledge. PTP-U formalises this and sets a new Pareto frontier on the forgetting–retention trade-off across eight strong baselines.
Figure 1: PTP-U separates "necessary" forgetting edits (Stage I: analytic proposal in K-FAC low-dim subspace) from "recoverable drift" on non-target capabilities (Stage II: projection back toward base policy). Achieves best forgetting–retention trade-off across 8 baselines.
Propose-Then-Project Unlearning (PTP-U) has two stages. Stage I applies closed-form K-FAC-curvature-guided analytic edits in a low-dimensional subspace, updating Attention then FFN with sequential re-linearisation between them, concentrating forgetting at the relevant parameters. Stage II projects back toward the original model's output distributions on non-target data under a nonlinear policy that keeps forgetting constraints satisfied. On MUSE-Books (Llama2-7B), PTP-U achieves BLEU 8.61 and ROUGE-L 6.33 with MMLU 43.34—the best combined forgetting and utility among 8 baselines (NPO_KL, RMU, WHP, ALTER, ASU, ICUL, SCANS, MET). On RWKU (Llama3.1-8B-Instruct), 94.20% average non-target utility with AA score 15.20 (lowest). On TOFU, ROUGE-L 0.10 (target) with GEK accuracy 0.79.
Figure 1: In the SusVibes benchmark, 56.9% of trajectories contain unfulfilled safety obligations vs. only 30.0% with forbidden actions—yet current guard models only flag the latter.
Guard models focused on forbidden actions miss the majority of agent safety failures. The authors introduce ObligationBench (240 expert-validated coding-agent trajectories) and ObligationGuard (Qwen3-8B fine-tuned on 40K GPT-5.6-Sol-generated synthetic trajectories), reaching 57.52% recall / 21.67% exact match vs. 48.97% / 10.00% for the best off-the-shelf model. Using ObligationGuard as feedback raises a Qwen3.8-27B coding agent's secure pass rate from 6.5% to 15.1% while barely affecting functional pass rate; ρ = 0.94 between ObligationBench recall and secure-pass improvement (arXiv:2610.11773).
Figure 1: IRIS exposes a critical gap in ASR-only safety evaluation: Toki Pona and CaesarCipher obfuscations both yield 2.9% ASR for GPT-4o, but operative understanding rates are 15.2% vs. 96.2%—meaning very different safety interpretations are warranted.
IRIS (Intent Recovery as an Intermediate Signal) adds an operative understanding rate (UR) alongside ASR: it measures whether a response both recognises the harmful task and treats it as the task to answer. Testing on 105 harmful tasks across 6 models, three increasingly explicit English reconstructions raise UR by 15.2–41.0 pp for all models while ASR shows no consistent direction—demonstrating that "safe" low-ASR responses can simply reflect obfuscation opacity rather than genuine alignment (arXiv:2610.11766).
Figure 1: Vanilla Ouro-1.4B ASR climbs from 29.9% at depth 1 to 61.8% at depth 3; DPO reduces d=2 but leaves deeper depths vulnerable. SafeBridge keeps ASR uniformly below 5% across all depths.
Looped LMs share weights across recurrent steps but show dramatically different jailbreak susceptibility by depth; vanilla Ouro-1.4B reaches 61.79% ASR at d=3. SafeBridge adds ~0.02% parameters via depth-specific FiLM modulation, selective state bridging, and safety-signed cross-depth supervision, reducing combined ASR from 21.06% (DPO) to 4.56% on 1.4B and from 44.93% to 9.85% on 2.6B; cross-depth transfer attacks drop from 66.24% to 12.02% (arXiv:2610.10625).
Figure 1: OnTrack compares each agent step's dependency graph against recorded successful runs using streaming optimal transport (~1ms/step). Using only 8 steps: +0.057 AUROC over baselines on SWE-bench; abort policy saves 18% compute.
OnTrack models agent trajectories as growing dependency graphs and compares them against reference successful runs via streaming OT (linear matching + quadratic structure penalty + KL penalties + entropy, α=0.45, β=0.20, δ=0.35). Using only the first 8 steps, it achieves +0.057 AUROC for failure prediction vs. content-similarity baselines on SWE-bench; an abort policy saves 18% of compute spent on failing runs with 83% of interruptions confirmed heading to failure (arXiv:2610.12375).
Figure 2: Held-out log-loss for 1D vs 3D IRT models on HarmBench. The 3D model (0.258) substantially outperforms the best 1D model (0.322) in all 5 splits, while DIF flags between OpenAI and Anthropic models drop from 13 to 1–2 under 3D scoring.
Applying a construct validity framework to HELM Safety, the authors find HarmBench fails to measure a single "harmful refusal" dimension: exploratory MIRT (1–10D) and confirmatory 3D/7D IRT models fit substantially better than 1D (log-loss 0.258 vs. 0.470), with multiple eigenvalues above the random-data cutoff. DIF analysis flags 13 OpenAI-vs-Anthropic items under aggregate scoring, collapsing to 1–2 under 3D scope-specific matching; the implication is that aggregate safety benchmarks conflate distinct harm behaviors and can inflate apparent developer differences (COLM 2026 AIMS Workshop, arXiv:2610.12409).
Figure 1: Small-world index (SWI) computed from attention-head activation similarity correlates with reasoning NLL across 6 LLMs. SWA-guided pruning at 0.7 sparsity reduces WikiText perplexity from 177.80 to 118.67 on Llama 3.2-3B with Wanda.
The authors build functional graphs from pairwise attention-head activation similarities and compute SWI (local clustering / path length, vs. random baseline), showing it tracks reasoning performance across 6 LLMs (Qwen3-14B SWI 1.82, NLL 0.082; Llama-1B SWI 0.92, NLL 0.169) and rises monotonically during training in Pythia-6.9B. Per-head core/bridge scores guide SWA sparsity allocation; at 0.7 sparsity with Wanda on Llama 3.2-3B, SWA reduces perplexity from 177.80 to 118.67 vs. 125.1 for ATP (arXiv:2610.12304).
Figure 1: RAG security catalog entity chain (Surface → Attack → Weakness → Risk → CIA Impact) and attack distribution across 38 attacks from 43 publications; 16 of 38 attacks target document ingestion, and 5 risks remain unmitigated.
A causal security meta-model for RAG systems, built from 43 publications (72% attack-focused), catalogs 3 RAG types, 8 surfaces, 38 attacks, 12 weaknesses, 41 risks, and 72 mitigations. Key findings: 42% of attacks target document ingestion (S-DOC), 5 risks are unmitigated (R17, R33–R36), output integrity has a 25% coverage gap, and in a worst-case multimodal/public/cloud deployment 40 of 41 risks remain applicable. Full OWASP LLM Top 10 coverage for 7 entries; interactive web visualizer released (arXiv:2610.11893).