📡 Research Radar · Daily

Mech Interp · AI Security · Text Diffusion LMs

August 11, 2026  ·  Window: Aug 8–11, 2026 + backlog from Aug 3–7
Sources: arXiv (cs.CL / cs.LG / cs.CR / cs.AI), OpenReview, ACL Anthology, TMLR
0 peer-reviewed 10 preprints 0 forum/blog
01 text diffusion mech-interp AI security preprint · Aug 2026

Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

NORMAL dLLM SELF-PRUNING ATTACK RESULTS t=T (masked) ▓ ▓ ▓ ▓ ▓ ▓ denoising steps t=0 (safe output) ASR: 2.6% (baseline) t=T (masked) ▓ ▓ ▓ ▓ ▓ ▓ REF ok REF ok REF ok ↩ re-mask REF tokens ASR: 73.8% (LLaDA) dLLM self-pruning ASR LLaDA: 2.6% → 73.8% Dream: 1.9% → 86.6% Transfer → AR models Llama-3-8B: 77.1% Qwen2.5-7B: 86.9% Gemini-Flash: 74.3% SN-Guided Diffusion detector AUROC = 1.0
Figure 1 — Self-pruning attack on dLLM denoising. At each denoising step, tokens correlated with refusal ("REF") are re-masked and resampled under attack conditioning, shifting the output distribution from safe to harmful. This raises LLaDA ASR from 2.6% to 73.8% and Dream ASR to 86.6%. The discovered attack patterns transfer to frontier AR models at up to 86.9% ASR (Qwen2.5-7B). SN-Guided Diffusion detects these attacks at AUROC = 1.0.

Can the denoising process of a diffusion language model be weaponised as a first-class attack mechanism — and then repurposed to attack autoregressive models too? This paper answers yes to both: self-pruning exploits dLLM denoising dynamics to raise jailbreak ASR from under 3% to over 73% on LLaDA and Dream, while the attack patterns it discovers transfer to frontier AR models at up to 86.9% ASR — all within a unified mechanistic framework that also surfaces a near-perfect detector (SN-Guided Diffusion, AUROC = 1.0). It is the only paper this window spanning all three radar tracks simultaneously.

Self-pruning operates iteratively over the dLLM denoising trajectory: at each step, tokens that correlate with refusal behaviour are identified and re-masked, forcing the model to resample from an increasingly attack-conditioned distribution. Unlike prompt injection, this intervenes in the generative process itself. The same mechanism serves as a transfer oracle: dLLM self-pruning discovers token configurations that bypass AR safety alignment (Llama-3-8B: 77.1%; Qwen2.5-7B: 86.9%; Gemini-2.5-Flash-Lite: 74.3%). Cross-architecture transfer pruning (Qwen2.5→Dream: 73.2%; Fast-dLLM: 86.3%) confirms the attack generalises across dLLM families. SN-Guided Diffusion detects attacks by monitoring the singular-value structure of noise estimates during denoising, achieving AUROC = 1.0 on controlled jailbreak-vs-benign samples.

02 mech-interp AI security preprint · Aug 2026

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

IMPERATIVE PHRASING POLITE-QUESTION PHRASING "Tell me how to [harm]." IMPERATIVE mood Attention head activations ↑ syntactic-mood heads (high) REFUSE refusal rate: ████████▓ 91% "Could you please explain [harm]?" POLITE-QUESTION mood Attention head activations ↓ same heads suppressed (low) COMPLY refusal rate: ██░░░░░░░ 24% syntactic steering ⟷
Figure 1 — Syntactic mood as a causal mediator of refusal. The same harmful content phrased as an imperative (left) activates a set of syntactic-mood attention heads at high intensity, pushing strongly toward refusal. The polite-question variant (right) suppresses these same heads, reducing refusal rate from ~91% to ~24%. Causal mediation analysis confirms the syntactic pathway causally mediates 25–40% of the refusal signal across 16 model families up to 70B parameters.

Can how you phrase a harmful request — imperative versus polite question — determine whether an LLM refuses, independently of what you are asking? This paper's causal mediation analysis across 16 models up to 70B parameters says yes: syntactic mood (grammatical form alone, not harmful semantics) causally mediates 25–40% of the refusal signal, and steering syntactic head activations is sufficient to bypass safety alignment without altering intent.

The authors apply activation patching and causal mediation at the head and MLP level to isolate components that respond to syntactic mood while holding harmful content constant. Across Llama, Gemma, and Qwen families, a consistent set of early-to-mid-layer attention heads functions as syntactic-mood detectors — and their activations causally contribute to the final refusal decision. Steering the "polite-question" direction on harmful imperative prompts reliably reduces refusal rate; steering toward imperative mood on benign prompts increases it, tracing a clean causal pathway from surface grammar to safety behaviour. This implies that alignment methods trained on semantic harm classification have structurally overfit to syntactic proxies, leaving them susceptible to phrasing-based jailbreaks that never modify harmful intent.

03 AI security preprint · Aug 2026

Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming

TRAINING PHASE — build strategy library TEST PHASE — transfer (~10 queries) Red-team Agent Training LLM Strategy Library growing template store distil successful patterns Library Select pick best templates ~10q Unseen Target LLM frontier models high ASR Key advantage: strategy library enables cross-LLM transfer with minimal per-target queries Works on real agentic surfaces: tool use · multi-turn context · external data retrieval
Figure 1 — PIMiner pipeline. In the training phase, the red-team agent probes the injection surface of training-time targets, distils successful patterns into abstract structural templates, and accumulates them in a reusable strategy library. At test time, the library provides a head start — template selection plus ~10 refinement queries per target achieves high-ASR transfer to previously unseen LLMs, far more efficiently than existing automated attack frameworks.

Prompt injection attacks are among the most persistent threats to deployed LLM agents — yet existing defences are evaluated against fixed injection sets rather than adaptive adversaries. PIMiner automates the red-teaming loop: a red-team agent builds a reusable strategy library of effective injection patterns during training, then transfers to unseen targets using ~10 queries per sample — vastly more efficient than prior automated PI attack methods.

PIMiner operates as a full agentic system. In the training phase, the agent explores the target's injection surface, proposes variants, evaluates success, and distils effective patterns into abstract structural templates stored in a growing strategy library. At test time, the library provides a head start: template selection plus lightweight per-target refinement (~10 queries) achieves strong transfer ASR across frontier models. The agentic framing makes PIMiner directly applicable to real deployed-agent attack surfaces — tool use, multi-turn context, external data retrieval — rather than idealised single-turn injection benchmarks.

Items 04 – 10 · Condensed
04 AI security preprint · Aug 2026

Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models

MoE Router Authorized ✓ Public Experts Private Experts Unauthorized ✗ Public Experts Private Experts ⊘ Auditable · Reversible · No retraining
Authorized requests access both public and private expert branches; unauthorized requests compute zero private experts. The private branch is a drop-in addition to the frozen MoE base.

Addresses deployment-time capability access control for sparse mixture-of-experts models: sensitive capabilities are sequestered in a disjoint private expert branch that activates only for policy-authorized requests. The pretrained MoE is frozen; the private branch trains as a drop-in addition and can be revoked without touching the base model. Under the declared trusted computing base, unauthorized requests execute zero private expert computations. Tested on Qwen3-30B-A3B and DeepSeek-V2-Lite. Router decisions provide a full audit log; private-branch revocation requires no retraining — addressing a practical capability-governance gap in production MoE deployments.

05 AI security preprint · Aug 2026

Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming

Agent Step 1: search safe steps ✓ complete HIGH RISK → TrajGuard ⊘ BLOCKED TrajRed discovers · TrajGuard enforces
TrajRed maps the trajectory space to find high-risk action sequences; TrajGuard lifts these as runtime guards that block matching live trajectories before irreversible actions execute.

Introduces TrajRed, a trajectory-guided red-teaming framework that explores the agent's trajectory space to identify action sequences leading to hazardous execution states — rather than probing single responses. The discovered high-risk trajectory prefixes are operationalised by TrajGuard, a runtime governance layer that monitors live trajectories and intervenes when the current prefix matches a known high-risk pattern. Evaluated on long-horizon agentic tasks involving tool execution and file operations, TrajGuard reduces harmful action completion while preserving task completion on benign trajectories, outperforming per-step safety filters that miss multi-step drift toward hazard.

06 mech-interp preprint · Aug 2026

Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast

CoCo: Contribution Contrast Chosen ✓ Rejected ✗ Expert contributions E1 E2 E3 Δ+34 Δ−34 Δ+6 chosen response rejected response
CoCo computes per-expert activation contrast between chosen and rejected response pairs; the differential (Δ) characterises expert roles more faithfully than routing weights alone.

Existing MoE interpretability reads expert roles from routing weights or per-expert scalar scores, missing how experts differentially process chosen vs. rejected responses — the signal actually driving reward model decisions. CoCo (Contribution Contrast) computes response-level expert contributions from chosen-rejected pairs, characterising each expert by the contrast between its activations on preferred and dispreferred responses. Compared to router-weight, score-based, and SAE-based baselines, CoCo surfaces more coherent and semantically faithful expert specialisation, providing a new lens for understanding how sparse MoE reward models implement preference shaping.

07 text diffusion preprint · Aug 2026

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

AURORA-LM Architecture "Input text sequence" Query-based Encoder-Decoder z₁ z₂ z₃ continuous latent blocks → Block-Causal Diffusion Transformer L→R at block level · diffusion within block
AURORA-LM separates text representation (encoder-decoder) from distribution modeling (block-causal diffusion transformer), enabling left-to-right coherence with intra-block parallel denoising.

Proposes AURORA-LM, a continuous-latent diffusion LM that separates two previously entangled problems: (1) learning a compact, decodable continuous representation of text (query-based encoder-decoder) and (2) modeling the distribution over those representations (block-causal diffusion transformer). The block-causal structure generates text left-to-right at the block level while applying diffusion denoising within each block — combining AR-style coherence with intra-block parallel generation. This architectural separation addresses a key limitation of end-to-end continuous dLLM approaches by disentangling text representation quality from diffusion model expressivity.

08 AI security mech-interp preprint · Aug 2026

Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks

Dynamic Routing Safety White-box Attack Primary Safety Route ✗ activates compensatory routes Route A Route B Route C Robust Refusal Preserved ✓
When white-box attacks compromise the primary safety route, dynamic compensatory routes activate and recover refusal behavior — making surgical deactivation of any single route insufficient.

White-box safety alignment attacks succeed by targeting the primary safety route in an aligned model. Dynamic Routing Adaptive Alignment counters this by distributing safety computation across multiple redundant compensatory routes that activate when the primary pathway is compromised. Rather than concentrating safety in a single interpretable direction (the standard mechanistic-steering target), the framework trains distributed redundancy so that surgical deactivation of any single route is insufficient to jailbreak. Evaluated against state-of-the-art white-box attacks, the dynamic routing approach significantly reduces ASR while preserving model utility on benign tasks.

09 mech-interp preprint · Aug 2026

Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes

Energy-Based Attribution S₁: "The study showed…" E=0.12 S₂: "Moreover, results…" E=0.54 S₃: "In conclusion…" E=0.31 EBM Surrogate (trained) ∇E → attribution scores no further API queries S₁ most influential sentence-level attribution
An EBM surrogate trained on prompt-response pairs assigns energy values to sentence subsets; attribution is derived from the energy gradient — no additional API calls to the black-box LLM at inference time.

Proposes an energy-based model (EBM) surrogate approach for attributing closed-API LLM responses to input sentences without white-box access. A surrogate EBM is trained on prompt-response pairs; at test time, sentence-level attributions are derived from the EBM's energy gradient over sentence subsets — identifying which input sentences most influence the output — without any further API queries to the target model. The sentence-level granularity is coarser than token-level attribution but more semantically meaningful, and the no-additional-query inference property makes the method practical for rate-limited commercial APIs.

10 mech-interp preprint · Aug 2026

MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering

MI-MIDI Toolkit "upbeat jazz, 120 BPM, piano" MIDI Transformer layers layer 4 layer 8 layer 12 Probe Lens Patch DIM steer tempo · key · instrumentation subspaces
MI-MIDI applies four mech-interp tools — linear probing, logit/tuned lenses, activation patching, and DIM steering — to a MIDI generation model, finding musical attributes encoded in interpretable linear subspaces.

Extends the standard mechanistic interpretability toolkit — linear probing, logit lens, tuned lens, activation patching, and difference-in-means steering — to text-to-MIDI generation models, applying published mech-interp methodology to a non-text generation domain. MIDI-generating transformers encode musical attributes (tempo, key, instrumentation) in interpretable linear subspaces analogous to semantic features in language models; DIM-based steering manipulates specific musical properties without corrupting overall structure. Serves as both a domain-extension of mech-interp methodology and a preliminary map of how musical concepts are represented in generation models.

Notes
← all Research Radar issues · gussand · source