Multiple authors — arXiv:2608.07430, August 7, 2026
Figure 1 — Self-pruning attack on dLLM denoising. At each denoising step, tokens correlated with refusal ("REF") are re-masked and resampled under attack conditioning, shifting the output distribution from safe to harmful. This raises LLaDA ASR from 2.6% to 73.8% and Dream ASR to 86.6%. The discovered attack patterns transfer to frontier AR models at up to 86.9% ASR (Qwen2.5-7B). SN-Guided Diffusion detects these attacks at AUROC = 1.0.
Can the denoising process of a diffusion language model be weaponised as a first-class attack mechanism — and then repurposed to attack autoregressive models too? This paper answers yes to both: self-pruning exploits dLLM denoising dynamics to raise jailbreak ASR from under 3% to over 73% on LLaDA and Dream, while the attack patterns it discovers transfer to frontier AR models at up to 86.9% ASR — all within a unified mechanistic framework that also surfaces a near-perfect detector (SN-Guided Diffusion, AUROC = 1.0). It is the only paper this window spanning all three radar tracks simultaneously.
Self-pruning operates iteratively over the dLLM denoising trajectory: at each step, tokens that correlate with refusal behaviour are identified and re-masked, forcing the model to resample from an increasingly attack-conditioned distribution. Unlike prompt injection, this intervenes in the generative process itself. The same mechanism serves as a transfer oracle: dLLM self-pruning discovers token configurations that bypass AR safety alignment (Llama-3-8B: 77.1%; Qwen2.5-7B: 86.9%; Gemini-2.5-Flash-Lite: 74.3%). Cross-architecture transfer pruning (Qwen2.5→Dream: 73.2%; Fast-dLLM: 86.3%) confirms the attack generalises across dLLM families. SN-Guided Diffusion detects attacks by monitoring the singular-value structure of noise estimates during denoising, achieving AUROC = 1.0 on controlled jailbreak-vs-benign samples.
Klerings, Brinkmann, Stuckenschmidt, Ponzetto — arXiv:2608.05409, August 5, 2026
Figure 1 — Syntactic mood as a causal mediator of refusal. The same harmful content phrased as an imperative (left) activates a set of syntactic-mood attention heads at high intensity, pushing strongly toward refusal. The polite-question variant (right) suppresses these same heads, reducing refusal rate from ~91% to ~24%. Causal mediation analysis confirms the syntactic pathway causally mediates 25–40% of the refusal signal across 16 model families up to 70B parameters.
Can how you phrase a harmful request — imperative versus polite question — determine whether an LLM refuses, independently of what you are asking? This paper's causal mediation analysis across 16 models up to 70B parameters says yes: syntactic mood (grammatical form alone, not harmful semantics) causally mediates 25–40% of the refusal signal, and steering syntactic head activations is sufficient to bypass safety alignment without altering intent.
The authors apply activation patching and causal mediation at the head and MLP level to isolate components that respond to syntactic mood while holding harmful content constant. Across Llama, Gemma, and Qwen families, a consistent set of early-to-mid-layer attention heads functions as syntactic-mood detectors — and their activations causally contribute to the final refusal decision. Steering the "polite-question" direction on harmful imperative prompts reliably reduces refusal rate; steering toward imperative mood on benign prompts increases it, tracing a clean causal pathway from surface grammar to safety behaviour. This implies that alignment methods trained on semantic harm classification have structurally overfit to syntactic proxies, leaving them susceptible to phrasing-based jailbreaks that never modify harmful intent.
Wang, Yin, Geng, Jia — arXiv:2608.05108, August 5, 2026
Figure 1 — PIMiner pipeline. In the training phase, the red-team agent probes the injection surface of training-time targets, distils successful patterns into abstract structural templates, and accumulates them in a reusable strategy library. At test time, the library provides a head start — template selection plus ~10 refinement queries per target achieves high-ASR transfer to previously unseen LLMs, far more efficiently than existing automated attack frameworks.
Prompt injection attacks are among the most persistent threats to deployed LLM agents — yet existing defences are evaluated against fixed injection sets rather than adaptive adversaries. PIMiner automates the red-teaming loop: a red-team agent builds a reusable strategy library of effective injection patterns during training, then transfers to unseen targets using ~10 queries per sample — vastly more efficient than prior automated PI attack methods.
PIMiner operates as a full agentic system. In the training phase, the agent explores the target's injection surface, proposes variants, evaluates success, and distils effective patterns into abstract structural templates stored in a growing strategy library. At test time, the library provides a head start: template selection plus lightweight per-target refinement (~10 queries) achieves strong transfer ASR across frontier models. The agentic framing makes PIMiner directly applicable to real deployed-agent attack surfaces — tool use, multi-turn context, external data retrieval — rather than idealised single-turn injection benchmarks.
Zhuoheng Huang, Mukesh Singh — arXiv:2608.06690, August 7, 2026
Authorized requests access both public and private expert branches; unauthorized requests compute zero private experts. The private branch is a drop-in addition to the frozen MoE base.
Addresses deployment-time capability access control for sparse mixture-of-experts models: sensitive capabilities are sequestered in a disjoint private expert branch that activates only for policy-authorized requests. The pretrained MoE is frozen; the private branch trains as a drop-in addition and can be revoked without touching the base model. Under the declared trusted computing base, unauthorized requests execute zero private expert computations. Tested on Qwen3-30B-A3B and DeepSeek-V2-Lite. Router decisions provide a full audit log; private-branch revocation requires no retraining — addressing a practical capability-governance gap in production MoE deployments.
Zhihao Zhu, Yi Yang (HKUST) — arXiv:2608.04018, August 4, 2026
TrajRed maps the trajectory space to find high-risk action sequences; TrajGuard lifts these as runtime guards that block matching live trajectories before irreversible actions execute.
Introduces TrajRed, a trajectory-guided red-teaming framework that explores the agent's trajectory space to identify action sequences leading to hazardous execution states — rather than probing single responses. The discovered high-risk trajectory prefixes are operationalised by TrajGuard, a runtime governance layer that monitors live trajectories and intervenes when the current prefix matches a known high-risk pattern. Evaluated on long-horizon agentic tasks involving tool execution and file operations, TrajGuard reduces harmful action completion while preserving task completion on benign trajectories, outperforming per-step safety filters that miss multi-step drift toward hazard.
Yifan Wang et al. — arXiv:2608.06400, announced August 2026
CoCo computes per-expert activation contrast between chosen and rejected response pairs; the differential (Δ) characterises expert roles more faithfully than routing weights alone.
Existing MoE interpretability reads expert roles from routing weights or per-expert scalar scores, missing how experts differentially process chosen vs. rejected responses — the signal actually driving reward model decisions. CoCo (Contribution Contrast) computes response-level expert contributions from chosen-rejected pairs, characterising each expert by the contrast between its activations on preferred and dispreferred responses. Compared to router-weight, score-based, and SAE-based baselines, CoCo surfaces more coherent and semantically faithful expert specialisation, providing a new lens for understanding how sparse MoE reward models implement preference shaping.
Multiple authors — arXiv:2608.02602, August 3, 2026
AURORA-LM separates text representation (encoder-decoder) from distribution modeling (block-causal diffusion transformer), enabling left-to-right coherence with intra-block parallel denoising.
Proposes AURORA-LM, a continuous-latent diffusion LM that separates two previously entangled problems: (1) learning a compact, decodable continuous representation of text (query-based encoder-decoder) and (2) modeling the distribution over those representations (block-causal diffusion transformer). The block-causal structure generates text left-to-right at the block level while applying diffusion denoising within each block — combining AR-style coherence with intra-block parallel generation. This architectural separation addresses a key limitation of end-to-end continuous dLLM approaches by disentangling text representation quality from diffusion model expressivity.
Multiple authors — arXiv:2608.02674, August 3, 2026
When white-box attacks compromise the primary safety route, dynamic compensatory routes activate and recover refusal behavior — making surgical deactivation of any single route insufficient.
White-box safety alignment attacks succeed by targeting the primary safety route in an aligned model. Dynamic Routing Adaptive Alignment counters this by distributing safety computation across multiple redundant compensatory routes that activate when the primary pathway is compromised. Rather than concentrating safety in a single interpretable direction (the standard mechanistic-steering target), the framework trains distributed redundancy so that surgical deactivation of any single route is insufficient to jailbreak. Evaluated against state-of-the-art white-box attacks, the dynamic routing approach significantly reduces ASR while preserving model utility on benign tasks.
Rezaee et al. (Sharif University of Technology) — arXiv:2608.02879, August 3, 2026
An EBM surrogate trained on prompt-response pairs assigns energy values to sentence subsets; attribution is derived from the energy gradient — no additional API calls to the black-box LLM at inference time.
Proposes an energy-based model (EBM) surrogate approach for attributing closed-API LLM responses to input sentences without white-box access. A surrogate EBM is trained on prompt-response pairs; at test time, sentence-level attributions are derived from the EBM's energy gradient over sentence subsets — identifying which input sentences most influence the output — without any further API queries to the target model. The sentence-level granularity is coarser than token-level attribution but more semantically meaningful, and the no-additional-query inference property makes the method practical for rate-limited commercial APIs.
Multiple authors — arXiv:2608.06638, August 7, 2026
MI-MIDI applies four mech-interp tools — linear probing, logit/tuned lenses, activation patching, and DIM steering — to a MIDI generation model, finding musical attributes encoded in interpretable linear subspaces.
Extends the standard mechanistic interpretability toolkit — linear probing, logit lens, tuned lens, activation patching, and difference-in-means steering — to text-to-MIDI generation models, applying published mech-interp methodology to a non-text generation domain. MIDI-generating transformers encode musical attributes (tempo, key, instrumentation) in interpretable linear subspaces analogous to semantic features in language models; DIM-based steering manipulates specific musical properties without corrupting overall structure. Serves as both a domain-extension of mech-interp methodology and a preliminary map of how musical concepts are represented in generation models.
Notes
All 10 items are preprints; no peer-reviewed papers in this window.
Item 01 (2608.07430) is the standout result: first paper to weaponise dLLM denoising mechanics as both a direct attack surface and a transfer-attack oracle against AR models, while simultaneously proposing a near-perfect detector — spanning all three radar tracks in a single paper.
Items 04, 05, 08 form an emerging cluster on architectural and agentic safety: capability access control in MoE (04), trajectory-guided red teaming (05), and distributed safety routing (08).
Items 04 (Policy-Masked Experts) and 06 (CoCo) together represent an emerging MoE-specific security and interpretability track — flag for weekly roundup alongside the dLLM cluster.
dLLM track: items 01 and 07 advance both the security (mechanistic exploits) and architecture (AURORA-LM continuous latent diffusion) dimensions.
Mech-interp methodology extensions this window: items 02 (syntactic mediation), 06 (MoE reward models), 09 (black-box energy attribution), 10 (music generation) each push the toolkit into a new domain.