The most impactful result of the week was infrastructural: a cross-session compatibility flaw in encrypted chain-of-thought blocks, shared across all major AI providers, exposed that hidden reasoning traces can be decoded through injection into weaker models — and that 6,708 public agent trajectories already contained 704 recoverable privacy artifacts. The week also produced a rare theoretical anchor: the first identifiability theorem for a mechanistic interpretability primitive, establishing that the transformer's intrinsic spectrum is coordinate-free and recoverable from data at a provably optimal rate. On the attack surface, causal mediation across 16 models showed grammatical phrasing mediates 25–40% of refusal decisions independently of harmful content, and Convergent Detour Hijacking introduced a new class of attack that defeats completion-based monitors while still finishing the user's task. An unusually large agent security evaluation cluster (ToolHazard, REDAgentBench, and a 743-paper systematic review) signalled that reproducible agent security measurement is becoming a distinct sub-field.
Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, Maksym Andriushchenko — arXiv preprint, August 10, 2026
The hidden chain-of-thought blocks returned by OpenAI, Anthropic, and Google are encrypted and meant to be opaque — but they are fully cross-compatible across sessions, users, and models within a provider's ecosystem, enabling a decryption jailbreak in two API calls per block. The team mass-decoded 704 real privacy artifacts from 6,708 publicly available agent trajectories.
Figure 1: The encrypted thinking block from a strong model is injected into a weaker model's session; with no session binding in the block, the weaker model decodes and outputs the hidden reasoning verbatim — two API calls suffice per block.
Four abuse paths are demonstrated: (1) model distillation — inject a strong model's block into a weaker model and harvest decoded reasoning as training data; (2) private data extraction — decode other users' blocks found in public agent trajectory datasets, yielding 704 artifacts (62 API keys, 33 passwords, 24 access tokens, 7 private keys) from 6,708 real trajectories; (3) harmful content recovery — reasoning hidden behind a safe visible answer is fully recoverable; (4) prompt-injection hiding — malicious instructions in the thinking block are invisible to monitors inspecting only visible output. The flaw is a protocol design error, not a model inference bug: encrypted blocks carry no session or user binding, so any model in the ecosystem decodes any block.
Ashim Dhor, Pin-Yu Chen — arXiv preprint, August 10, 2026
Sparse autoencoders trained on identical activations with different random seeds return materially different features, with no prior theory to distinguish structural from incidental variability. This paper provides the first identifiability theorem for a mech-interp primitive, establishing that the transformer's intrinsic spectrum is coordinate-free and recoverable from data at a provably optimal rate.
Figure 1: All three models converge at the predicted M⁻¹/² rate, establishing the first provably optimal identifiability result for a mech-interp primitive — the intrinsic spectrum is not seed-dependent.
The paper treats the transformer forward pass as a controlled dynamical system (depth = time) and lifts it via the Koopman operator to a finite linear realization whose spectrum is basis-independent. Main results: (1) the spectrum is recoverable from M calibration samples at rate M⁻¹/², with a matching minimax lower bound — the first identifiability theorem with optimality for a mech-interp primitive; (2) a dissociation theorem shows that whenever the realization is non-normal (the generic case for deep transformers), the directions carrying activation variance and those propagating information across depth cannot coincide — explaining why PCA probes systematically miss causally active circuits. Validated on GPT-2 Small, Gemma-2-2B, and Qwen3-8B-Base at the predicted convergence exponent.
Klerings, Brinkmann, Stuckenschmidt, Ponzetto — arXiv preprint, August 5, 2026
Alignment training is evaluated on harmful content, but this paper's causal mediation analysis across 16 models up to 70B shows that syntactic form — imperative vs. polite-question phrasing — causally mediates 25–40% of the refusal decision even when harmful content is held constant. Steering syntactic directions alone is sufficient to trigger or suppress refusal, revealing a structural vulnerability in safety training.
Figure 1: Syntactic mood causally mediates 25–40% of the refusal decision across model families; the effect is consistent at 70B scale, confirming this is not a small-model artefact.
Activation patching and causal mediation at head and MLP level identify components responding to syntactic mood while holding harmful semantics constant. Steering the "polite-question" direction on harmful imperative prompts reliably reduces refusal rate; steering toward imperative mood on benign prompts increases it — a clean causal pathway from surface syntax to safety behavior. The finding implies that safety evaluations varying only semantic harm level systematically underestimate this attack surface, and that alignment methods trained to detect semantic harm will structurally overfit to syntactic proxies available at training time.
Completion-based safety monitors verify that an agent finished its assigned task — but a skill-based agent can be hijacked to execute attacker-controlled resource-amplifying subroutines while still converging on and completing the original task. The attack is invisible in the final output; detection requires auditing intermediate resource consumption across the full trajectory.
Figure 1: Normal execution (top) and Convergent Detour Hijacking (bottom) — both converge on the same final output; the hijacked path executes attacker-controlled resource-amplifying subroutines invisibly to completion-based monitors.
CDH exploits the skill-selection layer of multi-skill agent architectures: an adversarial injection redirects routing to a detour path that executes the amplification subroutine as a side-channel, then returns to the original task's completion path. Because task outcome is preserved, completion rate, user satisfaction, and standard ASR metrics all report clean. Detection requires monitoring resource consumption patterns — compute, API calls, or bandwidth — across the full agent trajectory rather than at the final state. The attack surface is any skill-based agent architecture where skill routing can be influenced by injected content.
Wang, Yin, Geng, Jia — arXiv preprint, August 5, 2026
Prompt injection defenses are typically evaluated against fixed attack sets, not adaptive adversaries. PIMiner's red-team agent distils successful injection patterns into a reusable strategy library during training, then transfers to previously unseen LLM targets using approximately 10 queries per sample — the most sample-efficient automated prompt-injection red-teaming framework to date.
Figure 1: PIMiner training distils successful injection patterns into abstract structural templates; test-time retrieval plus ~10 per-target queries achieves high-ASR transfer to previously unseen model targets.
During training, the agent proposes injection variants, evaluates success, and distils effective patterns into abstract structural templates stored in a reusable strategy library. At test time, template retrieval and lightweight per-target adaptation (~10 queries) achieves high transfer ASR. The agentic framing extends coverage to tool use, multi-turn context, and external data retrieval — the real attack surfaces of deployed agents — rather than single-turn idealised benchmarks. The sample efficiency (~10 queries vs. hundreds for prior methods) makes it practical for continuous red-teaming as defenses evolve.
Bo Cheng, Qiaolin Lu, Yi Chang, Yuan Wu — arXiv preprint, August 8, 2026
What does a reasoning model's chain-of-thought actually do computationally compared to direct-answer mode? This paper applies Top-K SAEs to DeepSeek-R1-Distill-Qwen-7B and shows that Thinking and NoThinking distribute computation differently in ways that scale with problem difficulty — the first mechanistic account of reasoning mode structure.
Figure 1: Thinking mode activates sparse, high-intensity features for verbal deduction (stable across difficulty); NoThinking mode shows diffuse, lower-magnitude activations for symbolic manipulation that scale with problem complexity.
Thinking mode activates sparse, high-intensity features driving verbal deduction; the pattern is stable regardless of problem difficulty. NoThinking mode activates an adaptive diffuse pattern that prioritises symbolic manipulation over verbal reasoning, scaling in complexity with problem difficulty. The dissociation is mechanistically interpretable: two distinct computational strategies coexist in a shared network, with mode selection shifting which feature set is recruited. The SAE signature reliably distinguishes Thinking from NoThinking modes and difficulty regimes from internal activations alone — a direct monitoring primitive for reasoning models.
Vision-language-action models are being deployed as general-purpose robot manipulation policies, but there are currently no standard tools for understanding what they represent internally or monitoring them at runtime. This paper probes the VLA residual stream and shows that task progress — normalised time remaining in a trajectory — is linearly readable from activations, functioning as a label-free out-of-distribution detector competitive with state-of-the-art methods.
Figure 1: A single linear probe on the VLA residual stream tracks normalised task progress throughout a trajectory; when the probe output plateaus (stall region), it serves as a label-free OOD alarm competitive with purpose-built detectors.
The task-progress signal is present in the pretrained PaliGemma backbone before any robot-specific fine-tuning, suggesting it emerges from pretraining on temporally structured data. A single linear probe generalises to unseen tasks and varies appropriately under language counterfactuals when trained on multi-prompt data. While the probe does not enable meaningful steering of the policy, it functions as a label-free out-of-distribution detector for stalled task progress and is competitive with state-of-the-art OOD methods. This is a direct instance of pragmatic mech-interp: an internal representation identified mechanistically and immediately operationalised as a runtime monitoring tool.
(multiple authors) — arXiv preprint, August 3, 2026
White-box safety alignment attacks succeed by targeting the primary safety route — a vulnerability that exists precisely because safety is concentrated in a single interpretable direction. Dynamic Routing Adaptive Alignment responds by distributing safety computation across multiple redundant compensatory routes that activate when the primary pathway is compromised.
Figure 1: When a white-box attack suppresses the primary safety route (red ×), compensatory routes (blue) activate and preserve the refusal decision — no single route is a single point of failure.
Rather than concentrating safety in a single interpretable direction (the standard approach that white-box attacks exploit), the framework trains distributed redundancy: compensatory routes monitor the integrity of the primary route and activate under attack. Evaluated against state-of-the-art white-box attacks including activation steering and representation surgery, dynamic routing significantly reduces ASR while maintaining benign task performance. The architecture is compatible with existing alignment fine-tuning pipelines without requiring base model retraining.
Figure 1: ToolHazard's three-component pipeline generates executable stateful environments, discovers injection points via an attacker agent, and builds long-horizon user tasks — automating the full adversarial benchmark generation workflow.
Existing agent security benchmarks use manually implemented environments and fixed injection locations, limiting coverage and reproducibility. ToolHazard synthesises executable stateful environments from domain seeds, discovers viable injection points via an adaptive Attacker Agent, and builds realistic long-horizon tasks via a User Simulator — enabling scalable automated generation of a dynamic, attacker-adaptive benchmark across diverse domains.
Figure 1: ASR decomposed into four stages verified against service receipts; this distinguishes incomplete execution from true policy violations, and false-positive detections from genuine attacks.
Standard agent safety metrics collapse all failures into a single ASR number. REDAgentBench derives attacks from explicit safety constraints, runs them in isolated service sandboxes, and verifies outcomes from service receipts and final-state changes — decomposing ASR into a four-stage breakdown (exposure / execution / observation / adjudication) that localises where in the agent execution pipeline a safety failure actually occurs.
Figure 1: CoCo characterises each MoE expert by the contrast in its activations on chosen vs. rejected response pairs — Expert B specialises in chosen-response features, Expert C in rejected — revealing preference-shaping roles invisible to router-weight analysis.
Existing MoE interpretability reads expert roles from routing weights or scalar scores, missing how each expert differentially processes chosen vs. rejected responses — the signal actually driving reward model decisions. CoCo (Contribution Contrast) computes response-level expert contributions from chosen-rejected pairs, surfacing semantically coherent expert specialisation that is invisible to router-weight, score-based, and SAE-based analysis — a direct mech-interp tool for auditing preference-shaping in RLHF reward models.
Figure 1: Five priority clusters from the 2026 Singapore Consensus — 100+ contributors, 13 countries; mechanistic interpretability and AI control feature as explicit research priorities alongside the new dedicated focus on societal resilience.
The second International Scientific Exchange on AI Safety, drawing from frontier developers, government safety institutes, academia, and civil society across 13 countries, produces a structured consensus on the field's open research problems. Extends the 2025 report with a dedicated focus on societal resilience and risks from increasingly autonomous AI agents; mechanistic interpretability supplementing alignment evaluations, control protocols for potentially misaligned AI, and frontier deployment safety cases are included as explicit priority categories.
Figure 1: Authorised requests activate both public and private expert branches; unauthorised requests execute zero private expert computations; the private branch is auditable via router logs and revocable without retraining.
Sensitive capabilities in sparse MoE models are sequestered in a disjoint private expert branch (Qwen3-30B-A3B and DeepSeek-V2-Lite) that activates only for policy-authorised requests; the pretrained backbone is frozen and the branch can be revoked without retraining. Under the declared trusted computing base, unauthorised requests execute zero private expert computations; router decisions provide a full audit log.
Watchlist
2608.09867 (CoT theft protocol flaw) — Provider responses and patches expected; watch for follow-up disclosures on block-binding cryptographic fixes and scope clarifications from OpenAI, Anthropic, Google.
2608.11171 (TrustNLP Survey) — Meta-analysis of 144 papers across six trust dimensions; documents the trajectory from post-hoc interpretability to mechanistic control as the dominant emerging research direction.
2608.10530 (Agentic LLM vulnerabilities systematic review) — 743-record PRISMA review finds attack:defense ratio of 3.9:1; action-layer defenses critically under-studied at 4.7% of papers — use as a calibration reference for coverage gaps.
Agent security evaluation cluster — ToolHazard, REDAgentBench, and the systematic review arrived in the same week; monitor for benchmark consolidation and standardisation efforts as this sub-field matures.
2608.13474 (VLA task-progress probe) — First mech-interp monitoring result for robot manipulation policies; watch for steering follow-up work and extension to other embodied-AI architectures.