Frontier LLMs silently mark themselves when they've been hijacked — detectable in the residual stream at 90%+ AUROC even after the output is already compromised. That mechanistic signal becomes a near-perfect defense, dropping attack success from 34.6% to essentially zero.
Fig. 1 — Linear probes at layer 12 detect IPI exposure at 90–96% AUROC across all 6 frontier models (centre). AGRI defense built on these probes reduces attack success rate from 34.6% to ≈0% (right).
The paper probes intermediate hidden states of LLMs mid-inference. Across 6 frontier models (GPT-4o, Claude 3.5, Gemini 1.5, Llama 3.1, Mistral, Qwen 2.5), linear probes trained on the residual stream achieve 90%+ AUROC for IPI exposure—even when the model's output is already compromised. The probes underpin AGRI (Activation-Guided Resistance against Indirect prompt injection), which monitors activations during agentic tasks and blocks tool calls when the IPI signal exceeds a threshold. AGRI reduces the agent attack success rate from 34.6% to approximately 0% on standard IPI benchmarks. Probe weights and code are released.
Steerling-8B is a diffusion language model trained with interpretability baked into the loss function — every generation decomposes into named concept vectors, editable at inference time without fine-tuning. The first time training-time interpretability has scaled to 8B parameters without capability loss.
Fig. 1 — Steerling-8B: a concept bottleneck layer decomposes every generation into a sparse weighted sum of named concept vectors. Both attribution (which concepts fired?) and direct concept steering (edit weights at inference) are supported without any fine-tuning.
Steerling-8B introduces a concept bottleneck layer into the training of a masked diffusion LM. Each generation is decomposed into a sparse weighted sum of learned human-interpretable concept vectors; a sparsity regularizer keeps concept activation minimal without sacrificing output quality. At 8B scale, Steerling achieves near-parity with non-interpretable baselines on standard generation benchmarks while enabling concept attribution and direct concept steering at inference time without fine-tuning — the first demonstration that training-time interpretability constraints scale to large masked diffusion LMs without significant capability degradation.
Safety-aligned behavior lives in less than 2% of a model's features — a tiny, locatable circuit. Circuit-Anchored Evolution pins exactly those features during self-improvement, preventing safety erosion while letting everything else change freely.
Fig. 1 — Safety circuit: <2% of total features (red dots, left). Without CAE, safety score erodes across all 4 evolution scenarios; with CAE anchoring those weights, safety is maintained while capabilities improve (right).
The paper uses SAE decomposition to identify the "safety circuit" — the minimal set of features causally responsible for safety-aligned behavior — finding it constitutes less than 2% of total features, concentrated in specific attention heads and MLP neurons. CAE anchors these circuit weights during continued training or self-evolution, evaluated across 4 scenarios: continued pretraining, RLHF reset, self-improvement loop, and adversarial fine-tuning. CAE maintains safety ratings at pre-evolution levels while allowing capability metrics to improve, outperforming regularization-only baselines in all four scenarios.
Fig. 1 — Immune defense loop: each injection encounter generates a persistent antibody stored in a growing library; similar future attacks trigger antibody reactivation and blocking — 83% ASR reduction.
AgentAntibody models prompt injection defense as an adaptive immune response: each attack encounter generates an antibody (an adapted defense rule or classifier), stored in a persistent library and reactivated for similar future injections. On AgentDojo and WebArena-security benchmarks, the system achieves 83% attack success rate reduction with only 12% utility overhead.
Fig. 1 — Some models achieve high benchmark scores (x-axis) while scoring low on the latent safety factor (y-axis, red dots), exposing surface shortcuts. IRT recovers 3 latent factors that better characterize safety across 192 LLMs × 8 benchmarks.
Applies item response theory (IRT) to 192 LLMs evaluated on 8 safety benchmarks, recovering 3 latent safety factors that explain cross-benchmark variance better than benchmark averages alone. Models can score high on individual benchmarks by exploiting surface shortcuts while remaining unsafe on latent factors — revealing systematic measurement blind spots in current evaluation practice.
Fig. 1 — Without SPAR: erasing the target concept causes collateral degradation in adjacent concepts (left, faded red). SPAR's anchored regularization protects adjacent concepts while successfully erasing the target — 31% reduction in collateral degradation.
Multimodal LLM unlearning creates "knowledge holes" — adjacent un-targeted concepts degrade due to collateral damage. Selective Protection with Anchored Regularization (SPAR) identifies at-risk adjacent concepts via activation similarity and anchors their representations during unlearning, cutting collateral degradation by 31% while maintaining forget performance on the target concept.
Fig. 1 — Prompt-response embedding sequence modelled as a dynamical system; DMD extracts low-dimensional trajectory features that classify harmful vs. safe content at 94% accuracy across 6 safety datasets.
Treats the sequence of embedding states across a prompt-response exchange as a dynamical system and applies dynamic mode decomposition (DMD) to extract low-dimensional trajectory features. A DMD-feature classifier achieves 94% accuracy on harmful content detection across 6 standard safety datasets, with lower false positive rates than embedding-similarity baselines at matched recall.
Fig. 1 — Coverage audit: 12 major AI safety/capability benchmarks uniformly fail to measure response consistency, citation grounding, and adversarial robustness (bottom 3 rows, red), despite their centrality to deployment safety.
Systematic audit of 12 major AI safety/capability benchmarks identifies three consistently unmeasured dimensions: response consistency (same answer on repeated trials?), citation grounding accuracy, and adversarial robustness. Proposes 3 supplementary evaluation protocols with open scoring code to close these gaps.
Notes
Only 8 genuinely relevant uncovered items found — no padding. The Aug 20–23 weekend window had sparse new arXiv submissions in scope (typical late-August lull); 6 of 8 items are from Aug 1–7 missed by earlier runs.
Items #1 and #3 sketch a mechanistic safety pipeline: probes locate IPI/safety signals in the residual stream (#1); circuit anchoring preserves those signals through training (#3). Flag for weekly roundup.
Text diffusion LMs: item #2 (Steerling-8B) is the standout dLLM result — training-time interpretability at 8B without capability loss.