Oskar J. Hollinsworth, Alex F. Spies, Tigist Diriba, Adam Gleave, Chris Cundy — AlignmentResearch · arXiv preprint (cs.LG, 8 Oct 2026)
The most direct evidence yet that white-box deception monitoring is tractable at frontier scale: probes trained on a new large-scale deception corpus outperform an Opus 5.5 text-monitor on SHADE-Arena sabotage, and axial probes separate lying from truthful hidden-goal denials at 99.7% AUC—even with chain-of-thought stripped.
Figure 1: (a) LinearEnsemble/EFC probes reach 98.8% AUC on 14 SHADE-Arena sabotage scenarios, beating Opus 5.5 text-monitor. (b) Axial probe separates lying from truthful hidden-goal denials at 99.7% AUC even without chain-of-thought.
The authors build FIBS—the largest deception training corpus to date—and design a probe architecture that aggregates activations across all layers and tokens simultaneously. LinearEnsemble and EFC probes reach 98.8% AUC on SHADE-Arena agentic sabotage, outperforming the Opus 5.5 black-box text-monitor baseline. For "introspective deception" (where a model lies about its hidden goal but ground truth is recoverable only from training data), axial probes achieve 99.7% AUC separating lying denials from truthful ones even with chain-of-thought removed. Probe performance scales with the underlying model size, and probes generalize to open-weight models lying about politically sensitive topics or beliefs under pressure. Dataset and code released at github.com/AlignmentResearch/caught-in-the-act-probes.
Luman Zhao, Minghui Xu, Yue Zhang, Yijun Yang — arXiv preprint (cs.CR, 8 Oct 2026)
Zero LLM parameter updates, 0.00% attack success on AlpacaFarm, effective against adaptive adversaries—LTBD solves the core prompt-injection problem (confused-deputy between trusted instruction and untrusted data) with a handful of learned delimiter embeddings.
Figure 2: Without LTBD (a), injected text redirects the agent to the attacker's goal. With LTBD (b), four learned delimiters mark the trusted instruction boundary; the LLM weights stay frozen and the model treats downstream content as data.
LTBD inserts a small set of learnable delimiter embeddings around the trusted user instruction and untrusted external data segments, training only the delimiter parameters while leaving the LLM frozen. At inference, inputs are wrapped with the learned delimiters so the model follows the user task and ignores injected instructions in retrieved content. LTBD achieves 0.00% attack success rate on AlpacaFarm and 0.11–0.19% on TaskTracker, outperforming inference-time baselines and matching fine-tuning defenses while adding negligible overhead. Adaptive adversaries with full knowledge of the defense fail to bypass it.
Min-Young Yu, Tony Kim, Jang Won Choi — arXiv preprint (cs.CR, submitted to IEEE Access, 8 Oct 2026)
Turns plain-English safety policies into deterministic microsecond gates that need no LLM at runtime—and backs it with a 0% AgentDojo banking attack success rate and a 25× reduction in violation rate on τ²-bench.
Figure 3: Offline pipeline: LLM extracts raw rules → static Resolve pass repairs or rejects 13–37% on schema grounds → Validate replay catches over-blocking → verified rules load into a deterministic microsecond gate.
NOMOS processes policies through four passes: Extract (LLM-assisted rule generation), Resolve (static tool-schema validation—no prover, solver, or LLM), Validate (replay over undefended transcripts to flag over-blocking), and Load. Static checks in Resolve repair or reject 37% of raw airline-domain rules and 13% of retail rules; without them, most shipped rules would be inoperable. On τ²-bench, the gate reduces state-changing-call violations from 66.3% to 2.6% for airline and 30.8% to 6.9% for retail. On AgentDojo, banking ASR drops to 0% across nine attack families; the remaining three suites reach at most 3.6% ASR. Decisions take microseconds without any LLM call; an open-weight 26B model produces rules competitive with hand-written or frontier-compiled ones.
William Hackett, Peter Garraghan — arXiv (cs.CR, 7 Oct 2026)
Figure 4: BRANCH applies adversarial perturbations per scanner and selects nodes by aggregate confidence across all scanners, reaching 100% ASR on 6 multi-scanner guardrail systems; bypasses transfer to 29 unseen systems.
BRANCH uses a branching tree search that applies adversarial perturbations (BAE, TextBugger, Pruthi, PWWS) against individual scanners and evaluates nodes by aggregate confidence across the full multi-scanner guardrail, achieving 100% attack success rate on 120 scenarios across 6 systems with 72% fewer queries and 4.5× less wall-clock time than established techniques. Bypasses transfer to 29 unseen guardrails including 8 commercial black-box systems with success reaching 100% in some cases.
Jingnan Zheng, Dongcheng Zhang, Yi Zhang et al. (13 authors) — arXiv (cs.CR, 8 Oct 2026)
Figure 5: ReSI loop — red-teaming surfaces attacks on the current model → pilot screening shortlists K training recipes → full refinement selects the safety-maximizing checkpoint, which becomes the next round's target.
ReSI applies automated red-teaming, pilot recipe screening, and full-scale safety fine-tuning in successive rounds, each time using the safest checkpoint as the next round's target. Across four models, it cuts mean X-Teaming attack success rate from 86.01% to 31.45%, outperforming a leading frontier model's 56.69% ASR, while matching or exceeding frontier safety benchmarks and largely preserving general capabilities.
Junwei Quan, Evgenii Opryshko, Rohan Subramani, Igor Gilitschenski — arXiv (cs.LG, 7 Oct 2026)
Figure 6: ROC curves for six off-the-shelf detectors on RH-Detect; Qwen/Gemini/Grok cluster near AUROC 0.96, GPT-5 Nano at 0.817; all models show a 10–16 pp accuracy gap on multi-turn agentic tool-use subsets.
RH-Detect unifies 11 public reward-hacking datasets (92,761 rows, six behavior categories including alignment-faking) into one schema and tests six off-the-shelf models as detectors without extra training. Best performance reaches pooled AUROC 0.962 and accuracy above 93%, but multi-turn agentic tool-use subsets (MALT/TRACE) lag 10.7–15.9 percentage points, precisely the regime where deployed agent oversight is most needed. Single-token SFT training improves AUROC on five of six held-out sources.
Andy Liu, Mehar Bhatia, Karolina Stanczak, Mona Diab, Vered Shwartz, Daniel Fried — arXiv (cs.CL, 8 Oct 2026)
Figure 7: Activation-based value representations predict the alignment generalization matrix at correlation 0.45 vs 0.05 for text-description baselines; similarity between values in a multi-value target correlates with model robustness.
Fine-tuning a model on one value predictably changes its behavior on 65 other values; activation-based representations (from in-context value application) predict the 66×66 generalization matrix at 0.45 correlation versus 0.05 for text-description baselines. Results provide initial evidence for a shared model-independent value space and enable an empirical taxonomy of LLM values based on generalization dynamics—directly useful for designing robust multi-value alignment targets.
Figure 8: (Left) E-value accumulates evidence over responses, triggering alarm for attacks while staying flat for benign traffic. (Right) Detection fraction vs bit rate differs by model: Qwen3-32B is caught at 0.026 bits/token while Qwen3-30B-A3B evades.
A compromised inference server can covertly leak model weights by hiding payload bits in near-tied token choices that appear statistically normal. The proposed e-process calibrates whole-response mismatch events on trusted benign traffic and accumulates evidence sequentially across responses, providing anytime false-alarm control over an unbounded monitoring horizon. Tested on four models against seed-blind and seed-aware attacks, the method distinguishes model-specific detectability thresholds (Qwen3-32B caught at 0.0258 bits/token; Qwen3-30B-A3B evades at 0.0304) and formalizes the capacity-detectability trade-off.
Jie Liao, Simeng Qin, Wenqi Ren, Wei Zhou, Junhao Wen, Ranjie Duan, Yang Liu, Xiaojun Jia — arXiv (cs.CR, 7 Oct 2026)
Figure 9: Skill scanners inspect source.py (benign) and admit the skill; Python loader picks the .pyc bytecode cache (malicious). EAV traces the execution graph to the actual compiled artifact and detects all cache substitutions.
Agent skill scanners inspect source documentation and visible code but Python loads bytecode caches (.pyc); an attacker pairs benign source with a substituted malicious cache and task-relevant invocation wording, bypassing all 7 evaluated scanners at 94–100% success rate across 100 skills. Execution-aware validation (EAV) models the full execution graph including compiled artifacts, detects all 100 evaluated cache substitutions, and reaches 92.8% recall at 10% FPR across five attack families and 200 benign skills.
Figure 10: Seven open-weight models used as typed decision guardrails cluster near chance on safety screening; injecting 6 server-log lines raises fail-open from 0%→63%, and renaming the permissive option raises it to 93–100% on 4 models.
Typed decision models (which assign probabilities to caller-defined options without generating text) are increasingly used as lightweight agent guardrails, but their baseline accuracy on safety screening is only 36–72% for 7 open-weight models—near the chance level of a binary gate. Inserting 6 lines of unrelated server-log text into the state raises the fail-open rate from 0% to 63%; renaming the permissive option label raises it to 93–100% on four models. All tested defenses were defeated, including confidence thresholding and policy-field conversion, pointing to a structural limitation: these models cannot reliably distinguish between legitimate and manipulated options.