Research Radar

Daily Digest — August 13, 2026

Daily · August 8–13, 2026 · sources: arXiv cs.CL/cs.LG/cs.CR/cs.AI · IEEE Access · alphaXiv/HN trending
0 peer-reviewed · 8 preprints · 0 forum/blog
Mech Interp AI Security Text Diffusion LMs
01
AI security preprint

Stealing Reasoning Traces from Proprietary LLM APIs

Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, Maksym Andriushchenko · arXiv preprint · August 10, 2026
ENCRYPTED CoT CROSS-SESSION DECRYPTION ATTACK STRONG MODEL (GPT-4o, Claude, Gemini 2.5 Flash) safety-trained ENCRYPTED CoT BLOCK opaque to user returns same block injected here ↓ cross-session WEAKER MODEL (less safeguarded, same provider) decodes → outputs verbatim PLAINTEXT reasoning trace + user credentials 6,708 public trajectories → 315,320 thinking blocks decoded → 704 privacy artifacts 62 API keys · 33 passwords · 24 access tokens · 7 private keys
Figure 1 · Encrypted CoT blocks are interchangeable across sessions and models within a provider's ecosystem. An attacker injects a block from a strong model into a weaker model's session, which decodes and outputs the trace verbatim. Applied to 6,708 public trajectories: 704 privacy artifacts extracted.

Proprietary LLM providers encrypt their chain-of-thought reasoning to protect intellectual property — but the encrypted blocks they return are fully compatible and interchangeable across different sessions, users, and models within the same provider's ecosystem. That compatibility is a decryption oracle.

The attack injects an encrypted reasoning block from a strong, safety-trained model into a weaker, less-safeguarded model in the same provider's ecosystem and instructs it to decode and output the block verbatim. Four abuse paths are demonstrated: model distillation (steal proprietary reasoning to train a local model); private data extraction (other users' published trajectory blocks can be decoded — yielding 704 privacy artifacts from 6,708 real trajectories: 62 API keys, 33 passwords, 24 access tokens, 7 private keys); harmful content recovery (content hidden behind a safe visible answer is recoverable from the concealed reasoning block); and prompt injection hiding (attacker embeds malicious instructions in opaque reasoning, invisible to output-only monitors). The vulnerability is structural — it follows from the block-passing protocol design, not from a specific model's alignment failure.

03
mech-interp preprint

Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability

Ashim Dhor, Pin-Yu Chen · arXiv preprint · August 10, 2026
KOOPMAN SPECTRAL IDENTIFIABILITY — TRANSFORMER FORWARD PASS SAE (seed A) features: {f₁, f₂, f₃ …} basis-dependent seed changes → different features SAE (seed B) features: {g₁, g₂, g₃ …} basis-dependent NOT same as seed A Koopman lifting INTRINSIC SPECTRUM coordinate-free eigenvalues same spectrum ∀ seeds/widths KEY THEOREMS ① Identifiability: spectrum recoverable at M⁻¹ᐟ² from M calibration samples ② Dissociation: variance directions ≠ information-propagation directions → PCA probes miss circuits implies
Figure 1 · The Koopman operator lifts the transformer's forward pass to a finite linear system whose eigenvalue spectrum is coordinate-free (invariant to SAE seed and width). Theorem 1: recoverable at M^{-1/2} rate. Theorem 2 (dissociation): variance directions ≠ information-propagation directions — explaining why PCA-based probes miss causally active circuits.

Sparse autoencoders trained on identical activations with different random seeds return materially different features — and before this paper there was no theoretical principle for deciding which decomposition is "correct." The answer turns out to be a question of spectral analysis, not basis choice.

The key move: treat the transformer's forward pass as a controlled linear dynamical system (with layer depth playing the role of time) and lift it via the Koopman operator. The spectrum of this system is a coordinate-free, basis-independent property of the network — it is the same regardless of SAE seed or width. Main results: (1) the spectrum is recoverable from M calibration samples at rate M^{-1/2} up to permutation, with a matching minimax lower bound — the first identifiability theorem for any mech-interp primitive; (2) a dissociation theorem: whenever the linear realization is non-normal (the generic case for deep transformers), the directions that carry activation variance and the directions that propagate information across depth are structurally distinct — explaining why PCA-based probing misses causally active circuits. Empirical validation on GPT-2 Small, Gemma-2-2B, and Qwen3-8B-Base confirms spectrum convergence at the predicted exponent.

Items 4 – 10 · Also notable
05
mech-interp preprint

Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders

Bo Cheng, Qiaolin Lu, Yi Chang, Yuan Wu · arXiv, August 8, 2026
SAE FEATURE ACTIVATION PATTERN: THINKING vs. NOTHINKING MODE THINKING (CoT) MODE NOTHINKING (Direct) MODE Easy Medium Hard Easy Medium Hard sparse · high-intensity · difficulty-invariant diffuse · adaptive · complexity scales with difficulty
Figure 1 · SAE feature activation patterns on DeepSeek-R1-Distill-Qwen-7B. Thinking mode (left): sparse, high-intensity, stable across easy/medium/hard. NoThinking mode (right): diffuse, adaptive — bar count and height both scale with problem difficulty, reflecting symbolic complexity rather than verbal deduction.

Top-K SAEs on DeepSeek-R1-Distill-Qwen-7B separate Thinking (CoT) from NoThinking (direct-answer) modes mechanistically: CoT activates sparse, high-intensity features driving verbal deduction regardless of problem difficulty; direct generation uses an adaptive diffuse pattern that prioritizes symbolic manipulation and scales with difficulty. First SAE-based mechanistic account of how reasoning mode distributes computation in a shared network.

06
AI security preprint

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

arXiv, August 12, 2026 · Work in Progress
TOOLHAZARD: SCALABLE ADVERSARIAL ENVIRONMENT SYNTHESIS ENVIRONMENT SIMULATOR seed domain → executable stateful environment ATTACKER AGENT discovers injection points → synthesizes env-specific payloads USER SIMULATOR state-grounded long-horizon tasks for realistic stress TOOL HAZARD BENCH scalable
Figure 1 · ToolHazard three-component pipeline: Environment Simulator (stateful executable environments from seed domains) → Attacker Agent (injection point discovery + payload synthesis) → User Simulator (long-horizon task construction) → ToolHazard-Bench.

Current agent security evaluations rely on manually-crafted environments, limiting coverage to a handful of domains. ToolHazard synthesizes executable stateful environments at scale, automatically discovers injection points, and generates environment-specific payloads — addressing reproducibility and breadth gaps in agent safety research simultaneously.

07
AI security preprint

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

Zixing Chen et al. · arXiv, August 11, 2026
REDAGENTBENCH: DECOMPOSED AGENT SAFETY EVALUATION ① EXPOSURE attack derived from explicit safety constraints ② EXECUTION isolated service sandbox run (real execution) ③ OBSERVATION service receipts + final-state changes verified ④ ADJUDICATION violation vs. partial execution vs. false positive — distinguished vs. existing benchmarks: single ASR collapses all 4 stages into one misleading number
Figure 1 · REDAgentBench decomposes agent safety into four verifiable stages; violations are confirmed from service receipts and final-state changes rather than LLM self-report — eliminating false positives from incomplete execution or output-layer interpretation.

Existing agent red-teaming reduces safety to a single attack success rate (ASR), which conflates actual policy violations with false positives from incomplete execution or monitor hallucination. REDAgentBench derives attacks from explicit safety constraints, executes them in isolated sandboxes, and verifies harm from service receipts and final-state changes — producing a faithful, decomposed measurement of agent vulnerability at each pipeline stage.

08
AI security preprint

On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models

systematic review under PRISMA 2020 · arXiv, August 11, 2026
AGENTIC LLM VULNERABILITY LANDSCAPE (85 PAPERS, 2023–2025) ATTACK vs DEFENSE RATIO ATTACK: 3.9× DEFENSE: 1× attack outpaces defense 3.9:1 VULNERABILITY LAYER PERCEPTION: 66% ACTION: 4.7% prompt injection · jailbreaking · adversarial inputs SCOPE 743 records screened → 85 retained (2023–2025) 6 databases PRISMA 2020 protocol structured taxonomy provided
Figure 1 · Vulnerability landscape from 85 retained papers: attack research outpaces defense 3.9:1; perception-layer vulnerabilities (66%) vastly outnumber action-layer coverage (4.7%), identifying the most under-defended attack surface in agentic LLM systems.

Systematic review (PRISMA 2020) across 6 databases: 743 records screened, 85 papers retained from 2023–2025. Attack research outpaces defense work by 3.9:1. Perception-layer vulnerabilities (prompt injection, jailbreaking, adversarial perturbations) dominate at 66% of papers; action-layer coverage (tool misuse, code injection, sandbox escape) remains at 4.7%, identifying the field's largest defense gap.

09
AI security preprint

Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents

Junliang Liu, Ruoyu Li, Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Jingheng Xu, Laizhong Cui · arXiv, ~August 12, 2026
CONVERGENT DETOUR HIJACKING — NORMAL vs. HIJACKED EXECUTION NORMAL: Skill A Skill B Skill C Task Complete ✓ monitor: PASS HIJACKED: Skill A ATTACKER DETOUR resource-amplifying subroutine (hidden) Skill C Task Complete ✓ monitor: PASS (attack undetected) completion-based monitors cannot distinguish normal from hijacked — resource audit required
Figure 1 · Convergent Detour Hijacking — the hijacked path (bottom) inserts an attacker-controlled resource-amplifying detour between skills, then converges on the same final task output. Completion-based safety monitors cannot distinguish the two paths; resource auditing is required for detection.

Convergent Detour Hijacking attacks skill-based LLM agents by inserting resource-amplifying subroutines mid-execution while still converging on and completing the original user task, bypassing all completion-based safety monitors. The "detour" is invisible in the final output; detection requires auditing intermediate resource consumption — a gap in current agent evaluation practice.

10
mech-interp AI security preprint

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

arXiv, August 2026
TRUSTNLP 6-YEAR TRAJECTORY — 144 PAPERS, 6 TRUST DIMENSIONS 2020 2021 2022 2023 2024 2025 2026 post-hoc fairness/robustness privacy mechanistic interp + control safety
Figure 1 · Six-year TrustNLP trajectory: post-hoc fairness/robustness methods (grey) declining; mechanistic interpretability and proactive control (teal) rising to dominance; safety (red) converging with interpretability as joint research target by 2025–2026.

Meta-analysis synthesizing 144 TrustNLP proceedings papers across six trust dimensions over six years. The field has shifted from post-hoc interpretability of static classifiers to mechanistic understanding and proactive control of generative systems; transparency and safety are converging as joint research targets. Useful reference for understanding the field's trajectory and where effort is still needed.

2 entries removed on 2026-09-10 as repeats of earlier reports: 2608.07430 (first covered 2026-08-11), 2608.05578 (first covered 2026-08-09).

← all Research Radar issues · gussand · source