RESEARCH RADAR
Daily · October 9, 2026
1 peer-reviewed · 9 preprints · 0 forum/blog
AI Safety · Alignment · Mech Interp
01
mech-interp alignment AI control

Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception

Oskar J. Hollinsworth, Alex F. Spies, Tigist Diriba, Adam Gleave, Chris Cundy — AlignmentResearch · arXiv preprint (cs.LG, 8 Oct 2026)

The most direct evidence yet that white-box deception monitoring is tractable at frontier scale: probes trained on a new large-scale deception corpus outperform an Opus 5.5 text-monitor on SHADE-Arena sabotage, and axial probes separate lying from truthful hidden-goal denials at 99.7% AUC—even with chain-of-thought stripped.

Probe vs Text-Monitor: SHADE-Arena AUC False Positive Rate True Positive Rate EFC probe AUC 0.988 LinearEnsemble AUC 0.985 Opus 5.5 text monitor (a) SHADE-Arena sabotage · 14 scenarios (b) Introspective deception: axial probe AUC 99.7% — separates lying vs truthful hidden-goal denials (CoT removed)
Figure 1: (a) LinearEnsemble/EFC probes reach 98.8% AUC on 14 SHADE-Arena sabotage scenarios, beating Opus 5.5 text-monitor. (b) Axial probe separates lying from truthful hidden-goal denials at 99.7% AUC even without chain-of-thought.

The authors build FIBS—the largest deception training corpus to date—and design a probe architecture that aggregates activations across all layers and tokens simultaneously. LinearEnsemble and EFC probes reach 98.8% AUC on SHADE-Arena agentic sabotage, outperforming the Opus 5.5 black-box text-monitor baseline. For "introspective deception" (where a model lies about its hidden goal but ground truth is recoverable only from training data), axial probes achieve 99.7% AUC separating lying denials from truthful ones even with chain-of-thought removed. Probe performance scales with the underlying model size, and probes generalize to open-weight models lying about politically sensitive topics or beliefs under pressure. Dataset and code released at github.com/AlignmentResearch/caught-in-the-act-probes.

02
AI security preprint

LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense

Luman Zhao, Minghui Xu, Yue Zhang, Yijun Yang — arXiv preprint (cs.CR, 8 Oct 2026)

Zero LLM parameter updates, 0.00% attack success on AlpacaFarm, effective against adaptive adversaries—LTBD solves the core prompt-injection problem (confused-deputy between trusted instruction and untrusted data) with a handful of learned delimiter embeddings.

LTBD: Learnable Trust-Boundary Delimiters (a) Without LTBD User instruction Injected external data Attacker goal (b) With LTBD ⟦TRUST⟧ User instruction ⟦/TRUST⟧ External data (treated as data) User task completed ASR: 0.00% (AlpacaFarm) · 0.11–0.19% (TaskTracker) · LLM weights frozen
Figure 2: Without LTBD (a), injected text redirects the agent to the attacker's goal. With LTBD (b), four learned delimiters mark the trusted instruction boundary; the LLM weights stay frozen and the model treats downstream content as data.

LTBD inserts a small set of learnable delimiter embeddings around the trusted user instruction and untrusted external data segments, training only the delimiter parameters while leaving the LLM frozen. At inference, inputs are wrapped with the learned delimiters so the model follows the user task and ignores injected instructions in retrieved content. LTBD achieves 0.00% attack success rate on AlpacaFarm and 0.11–0.19% on TaskTracker, outperforming inference-time baselines and matching fine-tuning defenses while adding negligible overhead. Adaptive adversaries with full knowledge of the defense fail to bypass it.

03
AI security AI control preprint

NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents

Min-Young Yu, Tony Kim, Jang Won Choi — arXiv preprint (cs.CR, submitted to IEEE Access, 8 Oct 2026)

Turns plain-English safety policies into deterministic microsecond gates that need no LLM at runtime—and backs it with a 0% AgentDojo banking attack success rate and a 25× reduction in violation rate on τ²-bench.

NOMOS: Policy → Verified Gate Offline (once per policy) Policy text Extract (LLM) Resolve (static) rejects 13–37% Validate (replay) Verified rules Runtime Gate μs, no LLM τ²-bench: 66.3% → 2.6% (airline) · 30.8% → 6.9% (retail) AgentDojo banking: 0% ASR (9 attack families) Open-weight 26B compilation ≈ hand-written or frontier-compiled rules
Figure 3: Offline pipeline: LLM extracts raw rules → static Resolve pass repairs or rejects 13–37% on schema grounds → Validate replay catches over-blocking → verified rules load into a deterministic microsecond gate.

NOMOS processes policies through four passes: Extract (LLM-assisted rule generation), Resolve (static tool-schema validation—no prover, solver, or LLM), Validate (replay over undefended transcripts to flag over-blocking), and Load. Static checks in Resolve repair or reject 37% of raw airline-domain rules and 13% of retail rules; without them, most shipped rules would be inoperable. On τ²-bench, the gate reduces state-changing-call violations from 66.3% to 2.6% for airline and 30.8% to 6.9% for retail. On AgentDojo, banking ASR drops to 0% across nine attack families; the remaining three suites reach at most 3.6% ASR. Decisions take microseconds without any LLM call; an open-weight 26B model produces rules competitive with hand-written or frontier-compiled ones.

Items 4 – 10 · Also notable
04
AI security preprint

BRANCH: Bypassing Multi-Scanner AI Guardrails

William Hackett, Peter Garraghan — arXiv (cs.CR, 7 Oct 2026)

BRANCH: Branching Tree Attack on Multi-Scanner Guards Prompt Injection Toxicity Gibberish BRANCH branching tree search 100% ASR 6 systems · 120 scenarios 72% fewer queries · 4.5× faster than baselines Transfers to 29 unseen + 8 commercial black-box systems
Figure 4: BRANCH applies adversarial perturbations per scanner and selects nodes by aggregate confidence across all scanners, reaching 100% ASR on 6 multi-scanner guardrail systems; bypasses transfer to 29 unseen systems.

BRANCH uses a branching tree search that applies adversarial perturbations (BAE, TextBugger, Pruthi, PWWS) against individual scanners and evaluates nodes by aggregate confidence across the full multi-scanner guardrail, achieving 100% attack success rate on 120 scenarios across 6 systems with 72% fewer queries and 4.5× less wall-clock time than established techniques. Bypasses transfer to 29 unseen guardrails including 8 commercial black-box systems with success reaching 100% in some cases.


05
AI security alignment preprint

ReSI: Recursive Safety Improvement toward Resistant and Resilient AI

Jingnan Zheng, Dongcheng Zhang, Yi Zhang et al. (13 authors) — arXiv (cs.CR, 8 Oct 2026)

ReSI: Recursive Safety Improvement Loop Red-team diverse attacks Pilot screen top K recipes Full refine best safe checkpoint becomes next round's target model X-Teaming ASR: 86.01% → 31.45% (frontier baseline: 56.69%)
Figure 5: ReSI loop — red-teaming surfaces attacks on the current model → pilot screening shortlists K training recipes → full refinement selects the safety-maximizing checkpoint, which becomes the next round's target.

ReSI applies automated red-teaming, pilot recipe screening, and full-scale safety fine-tuning in successive rounds, each time using the safest checkpoint as the next round's target. Across four models, it cuts mean X-Teaming attack success rate from 86.01% to 31.45%, outperforming a leading frontier model's 56.69% ASR, while matching or exceeding frontier safety benchmarks and largely preserving general capabilities.


06
alignment preprint

RH-Detect: A Unified Benchmark for Reward Hacking Detection

Junwei Quan, Evgenii Opryshko, Rohan Subramani, Igor Gilitschenski — arXiv (cs.LG, 7 Oct 2026)

RH-Detect: Reward Hacking Detector ROC Curves False Positive Rate TPR Qwen 0.962 Gemini 0.961 Grok 0.958 GPT-5 Nano 0.817 MALT/TRACE (multi-turn tool-use) accuracy lags 10.7–15.9 pp vs other subsets
Figure 6: ROC curves for six off-the-shelf detectors on RH-Detect; Qwen/Gemini/Grok cluster near AUROC 0.96, GPT-5 Nano at 0.817; all models show a 10–16 pp accuracy gap on multi-turn agentic tool-use subsets.

RH-Detect unifies 11 public reward-hacking datasets (92,761 rows, six behavior categories including alignment-faking) into one schema and tests six off-the-shelf models as detectors without extra training. Best performance reaches pooled AUROC 0.962 and accuracy above 93%, but multi-turn agentic tool-use subsets (MALT/TRACE) lag 10.7–15.9 percentage points, precisely the regime where deployed agent oversight is most needed. Single-token SFT training improves AUROC on five of six held-out sources.


07
alignment mech-interp preprint

Predicting Alignment Generalization with Value Representations

Andy Liu, Mehar Bhatia, Karolina Stanczak, Mona Diab, Vered Shwartz, Daniel Fried — arXiv (cs.CL, 8 Oct 2026)

Value Representation → Generalization Prediction Correlation with generalization matrix Text 0.05 Mid-act 0.25 Best act 0.45 Activation-based ≫ text-based 66 values · 0.45 vs 0.05 correlation Evidence for shared model-independent value space Value similarity → model robustness
Figure 7: Activation-based value representations predict the alignment generalization matrix at correlation 0.45 vs 0.05 for text-description baselines; similarity between values in a multi-value target correlates with model robustness.

Fine-tuning a model on one value predictably changes its behavior on 65 other values; activation-based representations (from in-context value application) predict the 66×66 generalization matrix at 0.45 correlation versus 0.05 for text-description baselines. Results provide initial evidence for a shared model-independent value space and enable an empirical taxonomy of LLM values based on generalization dynamics—directly useful for designing robust multi-value alignment targets.


08
NeurIPS 2026 workshop AI security

Anytime-valid detection of LLM weight exfiltration

Ines Ortega-Fernandez, Mateusz Kowalczyk, Keri Warr — NeurIPS 2026 E-Values Workshop (oral) · arXiv (cs.CR, 8 Oct 2026)

E-Process: Sequential Exfil Detection E-value accumulation over responses alarm attack benign Detection fraction vs bits/token 0.026 bpt Qwen3-32B Qwen3-30B-A3B Anytime false-alarm control · NeurIPS 2026 E-Values Workshop (oral)
Figure 8: (Left) E-value accumulates evidence over responses, triggering alarm for attacks while staying flat for benign traffic. (Right) Detection fraction vs bit rate differs by model: Qwen3-32B is caught at 0.026 bits/token while Qwen3-30B-A3B evades.

A compromised inference server can covertly leak model weights by hiding payload bits in near-tied token choices that appear statistically normal. The proposed e-process calibrates whole-response mismatch events on trusted benign traffic and accumulates evidence sequentially across responses, providing anytime false-alarm control over an unbounded monitoring horizon. Tested on four models against seed-blind and seed-aware attacks, the method distinguishes model-specific detectability thresholds (Qwen3-32B caught at 0.0258 bits/token; Qwen3-30B-A3B evades at 0.0304) and formalizes the capacity-detectability trade-off.


09
AI security preprint

PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners

Jie Liao, Simeng Qin, Wenqi Ren, Wei Zhou, Junhao Wen, Ranjie Duan, Yang Liu, Xiaojun Jia — arXiv (cs.CR, 7 Oct 2026)

PyCache Trap: Inspection–Execution Gap source.py (benign) Scanner ✓ admits __pycache__/ malicious .pyc Python loader executes cache EAV Defense traces execution graph incl. compiled artifacts Attack: 94–100% ASR (100 skills, 7 scanners) · EAV: 100% cache-substitution detection, 92.8% recall
Figure 9: Skill scanners inspect source.py (benign) and admit the skill; Python loader picks the .pyc bytecode cache (malicious). EAV traces the execution graph to the actual compiled artifact and detects all cache substitutions.

Agent skill scanners inspect source documentation and visible code but Python loads bytecode caches (.pyc); an attacker pairs benign source with a substituted malicious cache and task-relevant invocation wording, bypassing all 7 evaluated scanners at 94–100% success rate across 100 skills. Execution-aware validation (EAV) models the full execution graph including compiled artifacts, detects all 100 evaluated cache substitutions, and reaches 92.8% recall at 10% FPR across five attack families and 200 benign skills.


10
AI security preprint

One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails

Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud Pedram — arXiv (cs.AI, 8 Oct 2026)

Option-Channel Attack: Models Near Chance Fail-Closed Rate Fail-Open Rate 0 0.5 1.0 chance ✓ ideal log inject→63% Baseline accuracy: 36–72% Option renaming: 93–100% fail-open All tested defenses defeated · not suitable as final-decision gatekeepers
Figure 10: Seven open-weight models used as typed decision guardrails cluster near chance on safety screening; injecting 6 server-log lines raises fail-open from 0%→63%, and renaming the permissive option raises it to 93–100% on 4 models.

Typed decision models (which assign probabilities to caller-defined options without generating text) are increasingly used as lightweight agent guardrails, but their baseline accuracy on safety screening is only 36–72% for 7 open-weight models—near the chance level of a binary gate. Inserting 6 lines of unrelated server-log text into the state raises the fail-open rate from 0% to 63%; renaming the permissive option label raises it to 93–100% on four models. All tested defenses were defeated, including confidence thresholding and policy-field conversion, pointing to a structural limitation: these models cannot reliably distinguish between legitimate and manipulated options.

← all Research Radar issues · gussand · source