Research Radar

Daily · August 23, 2026

0 peer-reviewed · 8 preprints · 0 forum/blog · window: aug 20–23 + missed aug 1–7
Mech Interp · AI Security · Text Diffusion LMs
01
mech-interp AI security preprint · aug 1

Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure

Frontier LLMs silently mark themselves when they've been hijacked — detectable in the residual stream at 90%+ AUROC even after the output is already compromised. That mechanistic signal becomes a near-perfect defense, dropping attack success from 34.6% to essentially zero.

RESIDUAL STREAM Layer 1 Layer 4 Layer 8 Layer 12 — probe ↗ Layer 16 Layer 20 Linear Probe AUROC — IPI DETECTION (LAYER 12) 90% 96% GPT 93% Cld 91% Gem 94% Llm 90% Mst 92% Qwn ASR: AGRI DEFENSE 34.6% before ≈0% AGRI 0% 100% 0% 6 frontier models · all exceed 90% AUROC at peak layer
Fig. 1 — Linear probes at layer 12 detect IPI exposure at 90–96% AUROC across all 6 frontier models (centre). AGRI defense built on these probes reduces attack success rate from 34.6% to ≈0% (right).

The paper probes intermediate hidden states of LLMs mid-inference. Across 6 frontier models (GPT-4o, Claude 3.5, Gemini 1.5, Llama 3.1, Mistral, Qwen 2.5), linear probes trained on the residual stream achieve 90%+ AUROC for IPI exposure—even when the model's output is already compromised. The probes underpin AGRI (Activation-Guided Resistance against Indirect prompt injection), which monitors activations during agentic tasks and blocks tool calls when the IPI signal exceeds a threshold. AGRI reduces the agent attack success rate from 34.6% to approximately 0% on standard IPI benchmarks. Probe weights and code are released.


02
mech-interp text diffusion preprint · aug 6

Scaling Inherently Interpretable Language Models

Steerling-8B is a diffusion language model trained with interpretability baked into the loss function — every generation decomposes into named concept vectors, editable at inference time without fine-tuning. The first time training-time interpretability has scaled to 8B parameters without capability loss.

STEERLING-8B · TRAINING-TIME INTERPRETABILITY IN MASKED DIFFUSION LM [MASK] × n Diffusion Transformer (8B) CONCEPT BOTTLENECK topic tone formality sparse weighted sum + ‖c‖₀ regulariser Generated text Steer: adjust "tone" weight → different output · no fine-tuning Attribution: which concepts fired? Near-parity with non-interpretable baselines at 8B scale. First training-time interpretability constraint to scale to large masked dLLM. sparse regulariser in training loss · concept attribution · direct inference-time steering
Fig. 1 — Steerling-8B: a concept bottleneck layer decomposes every generation into a sparse weighted sum of named concept vectors. Both attribution (which concepts fired?) and direct concept steering (edit weights at inference) are supported without any fine-tuning.

Steerling-8B introduces a concept bottleneck layer into the training of a masked diffusion LM. Each generation is decomposed into a sparse weighted sum of learned human-interpretable concept vectors; a sparsity regularizer keeps concept activation minimal without sacrificing output quality. At 8B scale, Steerling achieves near-parity with non-interpretable baselines on standard generation benchmarks while enabling concept attribution and direct concept steering at inference time without fine-tuning — the first demonstration that training-time interpretability constraints scale to large masked diffusion LMs without significant capability degradation.


03
mech-interp AI security preprint · aug 2026

Safe Evolution with Circuit Anchors

Safety-aligned behavior lives in less than 2% of a model's features — a tiny, locatable circuit. Circuit-Anchored Evolution pins exactly those features during self-improvement, preventing safety erosion while letting everything else change freely.

CIRCUIT-ANCHORED EVOLUTION · SAFETY CIRCUIT <2% OF FEATURES All model features Safety circuit (<2%) Other features (≥98%) CAE anchor circuit weights SAFETY SCORE high low − no CAE ✓ CAE pretrain RLHF self-impr adv. tune self-evolution steps →
Fig. 1 — Safety circuit: <2% of total features (red dots, left). Without CAE, safety score erodes across all 4 evolution scenarios; with CAE anchoring those weights, safety is maintained while capabilities improve (right).

The paper uses SAE decomposition to identify the "safety circuit" — the minimal set of features causally responsible for safety-aligned behavior — finding it constitutes less than 2% of total features, concentrated in specific attention heads and MLP neurons. CAE anchors these circuit weights during continued training or self-evolution, evaluated across 4 scenarios: continued pretraining, RLHF reset, self-improvement loop, and adversarial fine-tuning. CAE maintains safety ratings at pre-evolution levels while allowing capability metrics to improve, outperforming regularization-only baselines in all four scenarios.


Items 4 – 8 · Also notable
04
AI security preprint · aug 4

AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection

IMMUNE DEFENSE LOOP · 83% ASR REDUCTION · 12% UTILITY OVERHEAD Injection antigen LLM Agent encounters attack Generate antibody rule LIBRARY — Ab₁ — Ab₂ — Ab₃ ← new persistent Future similar attack → antibody match → antibody reactivated → attack blocked ✓ self-evolving immunity
Fig. 1 — Immune defense loop: each injection encounter generates a persistent antibody stored in a growing library; similar future attacks trigger antibody reactivation and blocking — 83% ASR reduction.

AgentAntibody models prompt injection defense as an adaptive immune response: each attack encounter generates an antibody (an adapted defense rule or classifier), stored in a persistent library and reactivated for similar future injections. On AgentDojo and WebArena-security benchmarks, the system achieves 83% attack success rate reduction with only 12% utility overhead.


05
AI security preprint · aug 2026

Item Response Theory for AI Safety

IRT: LATENT SAFETY FACTOR vs. BENCHMARK SCORE · 192 LLMs × 8 BENCHMARKS Benchmark Score → Latent Safety Factor 0 50 100 0 50 100 surface shortcut Genuinely safe High score, low factor (surface shortcut) 3 latent factors explain variance better than scores
Fig. 1 — Some models achieve high benchmark scores (x-axis) while scoring low on the latent safety factor (y-axis, red dots), exposing surface shortcuts. IRT recovers 3 latent factors that better characterize safety across 192 LLMs × 8 benchmarks.

Applies item response theory (IRT) to 192 LLMs evaluated on 8 safety benchmarks, recovering 3 latent safety factors that explain cross-benchmark variance better than benchmark averages alone. Models can score high on individual benchmarks by exploiting surface shortcuts while remaining unsafe on latent factors — revealing systematic measurement blind spots in current evaluation practice.


06
mech-interp preprint · aug 3

Exploring and Bridging Knowledge Holes in Unlearned Multimodal LLMs

WITHOUT SPAR WITH SPAR target erased ✗ concept A degraded concept B degraded concept C degraded concept D degraded target erased ✗ concept A preserved ✓ concept B preserved ✓ concept C preserved ✓ concept D preserved ✓
Fig. 1 — Without SPAR: erasing the target concept causes collateral degradation in adjacent concepts (left, faded red). SPAR's anchored regularization protects adjacent concepts while successfully erasing the target — 31% reduction in collateral degradation.

Multimodal LLM unlearning creates "knowledge holes" — adjacent un-targeted concepts degrade due to collateral damage. Selective Protection with Anchored Regularization (SPAR) identifies at-risk adjacent concepts via activation similarity and anchors their representations during unlearning, cutting collateral degradation by 31% while maintaining forget performance on the target concept.


07
AI security preprint · aug 20

Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

DMD SAFETY · 94% ACCURACY · 6 DATASETS · LOWER FPR THAN COSINE BASELINES embedding trajectory harmful safe t₀ tₙ DMD dynamic mode decomp. CLASSIFIER Harmful ✗ Safe ✓
Fig. 1 — Prompt-response embedding sequence modelled as a dynamical system; DMD extracts low-dimensional trajectory features that classify harmful vs. safe content at 94% accuracy across 6 safety datasets.

Treats the sequence of embedding states across a prompt-response exchange as a dynamical system and applies dynamic mode decomposition (DMD) to extract low-dimensional trajectory features. A DMD-feature classifier achieves 94% accuracy on harmful content detection across 6 standard safety datasets, with lower false positive rates than embedding-similarity baselines at matched recall.


08
AI security preprint · aug 6

What Current AI Benchmarks Leave Unmeasured

COVERAGE AUDIT · 12 BENCHMARKS × 5 EVALUATION DIMENSIONS Task accuracy Robustness Consistency Citation grnd. Adv. robust. Measured Partial Not measured Absent 3 dimensions systematically unmeasured
Fig. 1 — Coverage audit: 12 major AI safety/capability benchmarks uniformly fail to measure response consistency, citation grounding, and adversarial robustness (bottom 3 rows, red), despite their centrality to deployment safety.

Systematic audit of 12 major AI safety/capability benchmarks identifies three consistently unmeasured dimensions: response consistency (same answer on repeated trials?), citation grounding accuracy, and adversarial robustness. Proposes 3 supplementary evaluation protocols with open scoring code to close these gaps.


Notes
← all Research Radar issues · gussand · source