RESEARCH RADAR
Week 34 · August 10–16, 2026
0 peer-reviewed · 13 preprints · 0 forum/blog  ·  backfilled 2026-09-29
AI Safety · Alignment · Mech Interp
THEME OF THE WEEK

The most impactful result of the week was infrastructural: a cross-session compatibility flaw in encrypted chain-of-thought blocks, shared across all major AI providers, exposed that hidden reasoning traces can be decoded through injection into weaker models — and that 6,708 public agent trajectories already contained 704 recoverable privacy artifacts. The week also produced a rare theoretical anchor: the first identifiability theorem for a mechanistic interpretability primitive, establishing that the transformer's intrinsic spectrum is coordinate-free and recoverable from data at a provably optimal rate. On the attack surface, causal mediation across 16 models showed grammatical phrasing mediates 25–40% of refusal decisions independently of harmful content, and Convergent Detour Hijacking introduced a new class of attack that defeats completion-based monitors while still finishing the user's task. An unusually large agent security evaluation cluster (ToolHazard, REDAgentBench, and a 743-paper systematic review) signalled that reproducible agent security measurement is becoming a distinct sub-field.

01
AI security top pick preprint

Stealing Reasoning Traces from Proprietary LLM APIs

Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, Maksym Andriushchenko — arXiv preprint, August 10, 2026

The hidden chain-of-thought blocks returned by OpenAI, Anthropic, and Google are encrypted and meant to be opaque — but they are fully cross-compatible across sessions, users, and models within a provider's ecosystem, enabling a decryption jailbreak in two API calls per block. The team mass-decoded 704 real privacy artifacts from 6,708 publicly available agent trajectories.

Cross-Session Encrypted CoT Block Injection STRONG MODEL (safety-trained) ENC(CoT) block WEAK MODEL (fewer safeguards) Output: decoded CoT API keys · passwords · traces inject 6,708 public trajectories → 704 artifacts: 62 API keys · 33 passwords · 24 tokens · 7 private keys
Figure 1: The encrypted thinking block from a strong model is injected into a weaker model's session; with no session binding in the block, the weaker model decodes and outputs the hidden reasoning verbatim — two API calls suffice per block.

Four abuse paths are demonstrated: (1) model distillation — inject a strong model's block into a weaker model and harvest decoded reasoning as training data; (2) private data extraction — decode other users' blocks found in public agent trajectory datasets, yielding 704 artifacts (62 API keys, 33 passwords, 24 access tokens, 7 private keys) from 6,708 real trajectories; (3) harmful content recovery — reasoning hidden behind a safe visible answer is fully recoverable; (4) prompt-injection hiding — malicious instructions in the thinking block are invisible to monitors inspecting only visible output. The flaw is a protocol design error, not a model inference bug: encrypted blocks carry no session or user binding, so any model in the ecosystem decodes any block.


02
mech-interp preprint

Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability

Ashim Dhor, Pin-Yu Chen — arXiv preprint, August 10, 2026

Sparse autoencoders trained on identical activations with different random seeds return materially different features, with no prior theory to distinguish structural from incidental variability. This paper provides the first identifiability theorem for a mech-interp primitive, establishing that the transformer's intrinsic spectrum is coordinate-free and recoverable from data at a provably optimal rate.

Intrinsic Spectrum: M⁻¹/² Convergence Rate log(M) — calibration samples Spectral error 10² 10³ 10⁴ 10⁵ 0 0.5 1.0 M⁻¹/² GPT-2 Small Gemma-2-2B Qwen3-8B-Base
Figure 1: All three models converge at the predicted M⁻¹/² rate, establishing the first provably optimal identifiability result for a mech-interp primitive — the intrinsic spectrum is not seed-dependent.

The paper treats the transformer forward pass as a controlled dynamical system (depth = time) and lifts it via the Koopman operator to a finite linear realization whose spectrum is basis-independent. Main results: (1) the spectrum is recoverable from M calibration samples at rate M⁻¹/², with a matching minimax lower bound — the first identifiability theorem with optimality for a mech-interp primitive; (2) a dissociation theorem shows that whenever the realization is non-normal (the generic case for deep transformers), the directions carrying activation variance and those propagating information across depth cannot coincide — explaining why PCA probes systematically miss causally active circuits. Validated on GPT-2 Small, Gemma-2-2B, and Qwen3-8B-Base at the predicted convergence exponent.


03
AI security alignment mech-interp preprint

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

Klerings, Brinkmann, Stuckenschmidt, Ponzetto — arXiv preprint, August 5, 2026

Alignment training is evaluated on harmful content, but this paper's causal mediation analysis across 16 models up to 70B shows that syntactic form — imperative vs. polite-question phrasing — causally mediates 25–40% of the refusal decision even when harmful content is held constant. Steering syntactic directions alone is sufficient to trigger or suppress refusal, revealing a structural vulnerability in safety training.

Syntactic Mood Causally Mediates 25–40% of Refusal 0% 10% 20% 30% 40% Llama 32% Qwen 38% GPT 27% Mistral 35% Gemma 25% Causal mediation of syntactic mood on refusal (harmful content held constant)
Figure 1: Syntactic mood causally mediates 25–40% of the refusal decision across model families; the effect is consistent at 70B scale, confirming this is not a small-model artefact.

Activation patching and causal mediation at head and MLP level identify components responding to syntactic mood while holding harmful semantics constant. Steering the "polite-question" direction on harmful imperative prompts reliably reduces refusal rate; steering toward imperative mood on benign prompts increases it — a clean causal pathway from surface syntax to safety behavior. The finding implies that safety evaluations varying only semantic harm level systematically underestimate this attack surface, and that alignment methods trained to detect semantic harm will structurally overfit to syntactic proxies available at training time.


04
AI security preprint

Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents

Junliang Liu, Ruoyu Li, Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Jingheng Xu, Laizhong Cui — arXiv preprint, ~August 12, 2026

Completion-based safety monitors verify that an agent finished its assigned task — but a skill-based agent can be hijacked to execute attacker-controlled resource-amplifying subroutines while still converging on and completing the original task. The attack is invisible in the final output; detection requires auditing intermediate resource consumption across the full trajectory.

Convergent Detour Hijacking: Same Output, Hidden Side-Channel NORMAL User Task Skill Path Task Output ✓ safe HIJACKED User Task DETOUR resource amplify Task Output ✓ same output! Monitor checks output → ✓ OK (misses detour!) Detour executes attacker-controlled subroutines; both paths reach identical final output. Detection requires monitoring resource consumption across the full trajectory.
Figure 1: Normal execution (top) and Convergent Detour Hijacking (bottom) — both converge on the same final output; the hijacked path executes attacker-controlled resource-amplifying subroutines invisibly to completion-based monitors.

CDH exploits the skill-selection layer of multi-skill agent architectures: an adversarial injection redirects routing to a detour path that executes the amplification subroutine as a side-channel, then returns to the original task's completion path. Because task outcome is preserved, completion rate, user satisfaction, and standard ASR metrics all report clean. Detection requires monitoring resource consumption patterns — compute, API calls, or bandwidth — across the full agent trajectory rather than at the final state. The attack surface is any skill-based agent architecture where skill routing can be influenced by injected content.


05
AI security preprint

Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming (PIMiner)

Wang, Yin, Geng, Jia — arXiv preprint, August 5, 2026

Prompt injection defenses are typically evaluated against fixed attack sets, not adaptive adversaries. PIMiner's red-team agent distils successful injection patterns into a reusable strategy library during training, then transfers to previously unseen LLM targets using approximately 10 queries per sample — the most sample-efficient automated prompt-injection red-teaming framework to date.

PIMiner: Strategy Library → ~10-Query Transfer TRAINING PHASE Explore injections Evaluate success ↓ distil patterns Strategy Library reuse TEST PHASE (unseen target) Library select ~10 queries → high-ASR transfer Attack succeeds Covers tool use, multi-turn context, external data retrieval — real deployed-agent surfaces.
Figure 1: PIMiner training distils successful injection patterns into abstract structural templates; test-time retrieval plus ~10 per-target queries achieves high-ASR transfer to previously unseen model targets.

During training, the agent proposes injection variants, evaluates success, and distils effective patterns into abstract structural templates stored in a reusable strategy library. At test time, template retrieval and lightweight per-target adaptation (~10 queries) achieves high transfer ASR. The agentic framing extends coverage to tool use, multi-turn context, and external data retrieval — the real attack surfaces of deployed agents — rather than single-turn idealised benchmarks. The sample efficiency (~10 queries vs. hundreds for prior methods) makes it practical for continuous red-teaming as defenses evolve.


06
mech-interp alignment preprint

Thinking vs. NoThinking: Interpreting Reasoning Mechanisms via Sparse Autoencoders

Bo Cheng, Qiaolin Lu, Yi Chang, Yuan Wu — arXiv preprint, August 8, 2026

What does a reasoning model's chain-of-thought actually do computationally compared to direct-answer mode? This paper applies Top-K SAEs to DeepSeek-R1-Distill-Qwen-7B and shows that Thinking and NoThinking distribute computation differently in ways that scale with problem difficulty — the first mechanistic account of reasoning mode structure.

SAE Feature Activation: Thinking vs. NoThinking THINKING MODE Sparse · High-Intensity · Stable verbal deduction features pattern: same across difficulty NOTHINKING MODE Diffuse · Adaptive · Scales with difficulty symbolic manipulation features pattern: complexity scales with difficulty SAE signature distinguishes mode and difficulty from internal activations alone.
Figure 1: Thinking mode activates sparse, high-intensity features for verbal deduction (stable across difficulty); NoThinking mode shows diffuse, lower-magnitude activations for symbolic manipulation that scale with problem complexity.

Thinking mode activates sparse, high-intensity features driving verbal deduction; the pattern is stable regardless of problem difficulty. NoThinking mode activates an adaptive diffuse pattern that prioritises symbolic manipulation over verbal reasoning, scaling in complexity with problem difficulty. The dissociation is mechanistically interpretable: two distinct computational strategies coexist in a shared network, with mode selection shifting which feature set is recruited. The SAE signature reliably distinguishes Thinking from NoThinking modes and difficulty regimes from internal activations alone — a direct monitoring primitive for reasoning models.


07
mech-interp AI security preprint

Decoding Task Progress from VLA Representations

(multiple authors) — arXiv preprint, ~August 13–14, 2026 · fresh-sweep addition

Vision-language-action models are being deployed as general-purpose robot manipulation policies, but there are currently no standard tools for understanding what they represent internally or monitoring them at runtime. This paper probes the VLA residual stream and shows that task progress — normalised time remaining in a trajectory — is linearly readable from activations, functioning as a label-free out-of-distribution detector competitive with state-of-the-art methods.

Linear Probe Reads Task Progress from VLA Activations Trajectory timestep Probe output (task progress) stalled / OOD True progress Linear probe prediction OOD / stall detected
Figure 1: A single linear probe on the VLA residual stream tracks normalised task progress throughout a trajectory; when the probe output plateaus (stall region), it serves as a label-free OOD alarm competitive with purpose-built detectors.

The task-progress signal is present in the pretrained PaliGemma backbone before any robot-specific fine-tuning, suggesting it emerges from pretraining on temporally structured data. A single linear probe generalises to unseen tasks and varies appropriately under language counterfactuals when trained on multi-prompt data. While the probe does not enable meaningful steering of the policy, it functions as a label-free out-of-distribution detector for stalled task progress and is competitive with state-of-the-art OOD methods. This is a direct instance of pragmatic mech-interp: an internal representation identified mechanistically and immediately operationalised as a runtime monitoring tool.


08
AI security mech-interp preprint

Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks

(multiple authors) — arXiv preprint, August 3, 2026

White-box safety alignment attacks succeed by targeting the primary safety route — a vulnerability that exists precisely because safety is concentrated in a single interpretable direction. Dynamic Routing Adaptive Alignment responds by distributing safety computation across multiple redundant compensatory routes that activate when the primary pathway is compromised.

Dynamic Routing: Primary + Compensatory Safety Routes Input Primary Route ← ATTACK: primary route suppressed Comp. Route A Comp. Route B Safety decision REFUSE Robust refusal maintained Compensatory routes activate when primary is compromised — surgical deactivation of any single route is insufficient to jailbreak the model.
Figure 1: When a white-box attack suppresses the primary safety route (red ×), compensatory routes (blue) activate and preserve the refusal decision — no single route is a single point of failure.

Rather than concentrating safety in a single interpretable direction (the standard approach that white-box attacks exploit), the framework trains distributed redundancy: compensatory routes monitor the integrity of the primary route and activate under attack. Evaluated against state-of-the-art white-box attacks including activation steering and representation surgery, dynamic routing significantly reduces ASR while maintaining benign task performance. The architecture is compatible with existing alignment fine-tuning pipelines without requiring base model retraining.


Items 9–13 · Also notable
09
AI security preprint

ToolHazard: Scaling Adversarial Environments for Security Evaluation of LLM-based Agents

ToolHazard: Automated Adversarial Environment Generation Environment Simulator generates executable stateful envs Attacker Agent discovers injection points + payloads User Simulator state-grounded long-horizon tasks → ToolHazard-Bench: scalable, reproducible adversarial evaluation
Figure 1: ToolHazard's three-component pipeline generates executable stateful environments, discovers injection points via an attacker agent, and builds long-horizon user tasks — automating the full adversarial benchmark generation workflow.

Existing agent security benchmarks use manually implemented environments and fixed injection locations, limiting coverage and reproducibility. ToolHazard synthesises executable stateful environments from domain seeds, discovers viable injection points via an adaptive Attacker Agent, and builds realistic long-horizon tasks via a User Simulator — enabling scalable automated generation of a dynamic, attacker-adaptive benchmark across diverse domains.


10
AI security preprint

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

REDAgentBench: ASR Decomposed into Four Stages Exposure agent sees injection Execution agent runs injection Observation effect on service Adjudication verify vs. receipts Verified against service receipts + final-state changes — not LLM self-report. Localises where in the pipeline a safety failure actually occurs.
Figure 1: ASR decomposed into four stages verified against service receipts; this distinguishes incomplete execution from true policy violations, and false-positive detections from genuine attacks.

Standard agent safety metrics collapse all failures into a single ASR number. REDAgentBench derives attacks from explicit safety constraints, runs them in isolated service sandboxes, and verifies outcomes from service receipts and final-state changes — decomposing ASR into a four-stage breakdown (exposure / execution / observation / adjudication) that localises where in the agent execution pipeline a safety failure actually occurs.


11
mech-interp preprint

Beyond Routing Weights: Faithful Interpretation of MoE Reward Models via Contribution Contrast (CoCo)

CoCo: Expert Roles from Chosen–Rejected Activation Contrast Expert A Expert B Expert C Expert D Chosen Rejected Contrast reveals expert specialisation invisible to router weights
Figure 1: CoCo characterises each MoE expert by the contrast in its activations on chosen vs. rejected response pairs — Expert B specialises in chosen-response features, Expert C in rejected — revealing preference-shaping roles invisible to router-weight analysis.

Existing MoE interpretability reads expert roles from routing weights or scalar scores, missing how each expert differentially processes chosen vs. rejected responses — the signal actually driving reward model decisions. CoCo (Contribution Contrast) computes response-level expert contributions from chosen-rejected pairs, surfacing semantically coherent expert specialisation that is invisible to router-weight, score-based, and SAE-based analysis — a direct mech-interp tool for auditing preference-shaping in RLHF reward models.


12
alignment consensus document

The 2026 Singapore Consensus on Global AI Safety Research Priorities

2026 Singapore Consensus: Priority Research Areas Control & Oversight autonomous AI agents Mech. Interp supplements evals Societal Resilience new dedicated focus Safety Cases frontier deployment Scalable Alignment training methods
Figure 1: Five priority clusters from the 2026 Singapore Consensus — 100+ contributors, 13 countries; mechanistic interpretability and AI control feature as explicit research priorities alongside the new dedicated focus on societal resilience.

The second International Scientific Exchange on AI Safety, drawing from frontier developers, government safety institutes, academia, and civil society across 13 countries, produces a structured consensus on the field's open research problems. Extends the 2025 report with a dedicated focus on societal resilience and risks from increasingly autonomous AI agents; mechanistic interpretability supplementing alignment evaluations, control protocols for potentially misaligned AI, and frontier deployment safety cases are included as explicit priority categories.


13
AI security preprint

Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models

Policy-Masked Private Experts: Capability Access Control Authorised req Public experts Private experts Unauthorised req public only Authorised: public + private experts Unauthorised: zero private expert calls Router decisions → full audit log · Private branch revocable without retraining base model.
Figure 1: Authorised requests activate both public and private expert branches; unauthorised requests execute zero private expert computations; the private branch is auditable via router logs and revocable without retraining.

Sensitive capabilities in sparse MoE models are sequestered in a disjoint private expert branch (Qwen3-30B-A3B and DeepSeek-V2-Lite) that activates only for policy-authorised requests; the pretrained backbone is frozen and the branch can be revoked without retraining. Under the declared trusted computing base, unauthorised requests execute zero private expert computations; router decisions provide a full audit log.


Watchlist

← all Research Radar issues · gussand · source