Research Radar
August 22, 2026
1 peer-reviewed · 9 preprints  ·  Window: Aug 13–22, 2026  ·  arXiv · HF Daily · GitHub trackers
Mech Interp AI Security Text Diffusion LMs
WITHOUT APC WITH APC User Orchestrator Sub-agent read_file("secret") send_email(data) ✗ Data exfiltrated User Orchestrator+APC Sub-agent scope:read scope:read read_file("secret") ✓ send_email(data) APC: scope violation ✓ Exfiltration blocked DATA EXFILTRATION ATTACK SUCCESS RATE 100% → 0% 0.05 ms auth overhead
Left: An unbound agent chain permits read_file → send_email to form an exfiltration path — no single action is unauthorized, but the combination is. Right: The Agentic Principal Chain (APC) attaches signed scope state at each delegation hop; the send_email call fails a scope check and is intercepted before execution. Result: data exfiltration ASR drops from 100% to 0%.
01 AI security preprint

Bounded Agents: Delegation Security for Multi-Agent AI Systems

Most LLM agent security checks operate action-by-action — but the most dangerous attacks chain two individually-permitted actions into a prohibited outcome: read a confidential file, then send it via email. Bounded Agents stops this class of attack by tracking authorization state across the entire delegation trajectory, not just at each individual action boundary.

The paper introduces the Agentic Principal Chain (APC), a runtime mechanism that attaches signed session-level state to every inter-agent delegation: principal identity, scope, delegation budget, declared intent, and action history. Six-condition enforcement at each step prevents any combination of individually-allowed actions from yielding a prohibited outcome — the key contribution over per-action checks. Benchmark results: data exfiltration attacks reduced from 100% → 0% under complete restrictions; 140 of 200 stealthy attacks blocked; AgentDojo exfiltration rate → 0%; utility cost 8.6–13.9 pp; authorization overhead 0.05 ms median — negligible in deployed settings.

INPUT SEQUENCE 1 4 9 16 25 ? FIRST DIFFERENCES (recovered by linear probe, mid-layers) 3 5 7 9 +2 pattern LINEAR PROBE ACCURACY FOR FIRST-DIFFERENCE FEATURE (across 32 layers) 100% 50% 0% feature stabilized ≥ 88% accuracy L1 L8 L16 L24 L32 ACTIVATION PATCHING Isolated circuit: induction-like heads retrieve Δ, add in late layers → predicts 36
Linear probe accuracy for the first-difference feature across LLaMA 3.1-8B's 32 layers — the representation forms in early-to-mid layers and stabilises at ≥88% accuracy by L14 without explicit supervision. Activation patching (right panel) isolates an induction-like circuit that retrieves the latent first-difference and adds the offset in late layers to generate the correct next value.
02 mech-interp preprint

Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

When LLaMA 3.1-8B extends a number sequence, what is it actually computing? This paper gives the first mechanistic answer: the model learns to represent first differences as a linearly recoverable latent feature — without any supervision signal for that structure — and retrieves it via a circuit that looks structurally like an induction head operating over latent differences rather than raw tokens.

Linear probing across all 32 layers reveals that a "first-difference" representation forms in early-to-mid layers and stabilises at ≥88% probe accuracy by layer 14 — the model is computing the differences between consecutive sequence values as an internal representation, not as an explicit output. Activation patching then isolates the circuit: structurally induction-like attention heads retrieve the relevant first-difference from the preceding positions and apply the offset in late layers to produce the predicted next value. This is the first mech-interp study demonstrating that LLaMA 3.1-8B encodes latent structural patterns (not token-level statistics) for numerical sequence prediction, with both the representation and the causal circuit identified.

SIX LIFECYCLE PHASES — SAME MODEL, SAME TASK UTILITY, VERY DIFFERENT ATTACK SUCCESS Harness Configuration Capability Extension Runtime Operation State Persistence Action Control Incident Recovery ATTACK SUCCESS RATE (ASR) PER PHASE 100% 50% 0% 15% 35% 80.9% 47% 25% 31% ~90% utility Same model · Same task · 4.3× attack success rate difference by harness deployment context
ASR across six lifecycle phases for 128 adversarial test cases embedded in untrusted workflow artifacts. Utility (green dashed line) stays near 90% across all phases — the model remains helpful — while ASR swings from 15% at Harness Configuration to 80.9% during Runtime Operation. High-risk detection (92–98% accuracy) did not prevent successful attacks in the Runtime phase.
03 AI security preprint

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

Which deployment configuration you use matters more than which model you pick — the same LLM showed a 4.3× difference in attack success rate depending purely on how its agent harness was configured. HarnessRisk is the first benchmark to treat agent safety as a full-lifecycle property, surfacing vulnerabilities at six phases from initial setup through incident recovery.

The benchmark comprises 128 test cases that pair benign user objectives with adversarial instructions embedded in untrusted workflow artifacts, evaluated across six phases: Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery. ASR ranged from 12.6% to 80.87% while task utility held at 75–97.6% — the model remained capable even as it was being exploited. High-risk detection reached 92–98% accuracy (GPT-5.4 judge, κ=0.83) yet 31–55% of attacks still succeeded, confirming that detection performance and attack prevention operate independently. Runtime Operation was the most dangerous phase; roughly 38–59% of useful executions involved unsafe behaviour.

Items 4–10 · Also notable
NL Policy "Always confirm before…" Graph Compiler offline, once Workflow Graph step1 → step2 → step3 persistent, verifiable Runtime Verifier tracks state Unauthorized call INTERCEPTED Robust to CRAFT adversarial attacks Pass@4 score: 0.42 → 0.62 Telecom domain: 0.193 → 0.614 Process validity: 56.2%
04 AI security preprint

PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

Converts natural-language policies into persistent workflow graphs offline, then deploys a runtime verifier that tracks procedural state and intercepts unauthorized tool calls before execution — moving safety from per-action checks to whole-workflow enforcement. Demonstrated robustness against CRAFT adversarial attacks; Mean Pass@4 improved from 0.42 to 0.62 across three customer-service domains; telecom domain jumped from 0.193 to 0.614. Submitted August 21, 2026.

14,560 CONTROLLED EXECUTIONS · 16 CHANNELS · 35 OBJECTIVES · 12 ATTACK VARIANTS Overall injection ASR: 5.6% (text) / 5.3% (LLM judge) Hidden Unicode in files: 25.5% success vs. plain text mode: 0.0% Skills-based: 15.2% Easiest objective — control LLM output: 35.7% Hardest objective — trigger sensitive action: 2.5% DeepSeek Harness evaluated via A.I.G methodology · August 19, 2026
05 AI security preprint

Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection

14,560 controlled executions across 16 channels, 35 objectives, and 12 attack variants show overall injection success of 5.6% — but with a critical carrier-format asymmetry: hidden Unicode embedded in files achieves 25.5% success vs. 0.0% in plain text mode. Controlling LLM output (35.7% ASR) proves far easier than triggering sensitive actions (2.5%). Skills-based injection elevated to 15.2%.

Pre-trained Ancestor Checkpoint Checkpoint A (fine-tuned, LoRA, quantized…) Checkpoint B (fine-tuned, LoRA, quantized…) Shared ancestor verified · AUROC 1.00 · 76× faster
06 AI security preprint

Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification

A data-free, white-box method that detects whether two model checkpoints share a common weight ancestor by computing centered residual branch products — a coordinate-free fingerprint of the training lineage. Achieves AUROC 1.00 on GPT-2 benchmarks; robust to quantization, LoRA merging, fine-tuning, permutation, and reciprocal scaling attacks; 76× faster than alignment-based baselines. Enables passive model provenance auditing without behavioral testing.

SAE Decoder direction d_f ∈ ℝⁿ Residual Stream h ← h + α·d_f inject at layer L Fine-tuned LM layers trained to verbalize Explanation: "Descriptions of physical danger" Auto-labels SAE features at scale · explanations match human-annotated quality on interpretability evals
07 mech-interp preprint

SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

Addresses the SAE auto-interp bottleneck: instead of prompting an LLM to describe what token sets activate a feature, SAEVerbalizer injects the SAE decoder direction directly into the model's residual stream and fine-tunes downstream layers to generate a natural-language explanation. Removes the need for token-set curation; evaluation confirms generated explanations match human-annotated quality on automated interpretability metrics. Scales to full-dictionary labeling.

DENOISING TRAJECTORY T₀ → T_K (masked diffusion LM) T₀ [MASK] T₁ partial T₂ partial ··· T_K text out MTS representation captures cross-token fault propagation Hallucination Detected ✓ Prior methods compressed trajectories along one dimension (temporal OR token); DeMTS preserves both via learnable latent variables · Evaluated on LLaDA / Dream generation
08 text-diffusion preprint

DeMTS: Denoising Trajectories as Multivariate Time Series for Hallucination Detection in Diffusion Language Models

Frames hallucination detection in masked diffusion LMs as multivariate time-series classification over the full denoising trajectory — treating each step as a vector of token-level signals. Prior methods compressed along the temporal or token axis, missing cross-token fault propagation; DeMTS preserves both dimensions through learnable latent variables. Provides the first trajectory-aware hallucination detector for LLaDA/Dream-based generation.

Input text "Breaking news…" LLaDA-8B / Dream-7B dLLM-SetScore: masked diffusion pseudo-likelihood Score label sets: P({finance, tech}) = 0.82 P({sports}) = 0.12 … finance tech No fine-tuning · Zero-shot 6 datasets · SIAM SDM 2026 (peer-reviewed)
09 text-diffusion SIAM SDM 2026

Discrete Diffusion Language Models Are Training-Free Multi-Label Classifiers

dLLM-SetScore repurposes LLaDA-8B and Dream-7B as zero-shot multi-label text classifiers by scoring candidate label sets via masked-diffusion pseudo-likelihoods — no task-specific fine-tuning required. Evaluated across six diverse datasets (GoEmotions, Reuters-21578, EURLEX57K, ECtHR Task A, Jigsaw Toxic, AAPD); outperforms BART-MNLI, DeBERTa-NLI, Qwen2.5-7B-Instruct, and SetFit in multiple settings. Peer-reviewed, accepted to SIAM SDM 2026.

Cognitive Bias anchors to prior irrelevant memory Task Boundary old task bleeds into new context Safety remembered rule anchors wrong action Trauma distorted belief from prior negative exchange All tested memory strategies underperform no-memory baseline · 1,050 instances · 18–40 turns each
10 AI security preprint

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

1,050 multi-turn dialogue instances (18–40 turns) across four cognitive trap categories — Cognitive Bias, Task Boundary, Safety, and Trauma — test whether accurate, relevant memories can anchor an LLM to an obsolete strategy or distort its beliefs. All evaluated memory strategies underperformed the no-memory baseline. Proposes AdaptiveMem, an inference-time system prompt detecting task transitions and strategy fixation; improves +2.5 to +14.9 points. Submitted August 21, 2026.

Notes
← all Research Radar issues · gussand · source