Research Radar
Week 32 · July 27 – August 2, 2026 (backfilled 2026-09-29)
2 peer-reviewed · 12 preprints · 0 forum/blog
AI Safety · Alignment · Mech Interp
Theme of the Week

This week's dominant signal was a methodological audit of sparse autoencoder interpretability. Three independent papers converged on the finding that SAE evaluations systematically confuse distinct axes: geometric recovery is not causal function, activation descriptions do not predict downstream effect, and pipeline configuration explains more variance in autointerpretability scores than SAE architecture does. ParityTransformer raised the ante by demonstrating that training-time interpretability constraints outperform post-hoc SAEs on every causal metric, pointing toward a possible paradigm shift away from the retrofit programme. On the security side, the week's most striking pattern was the proliferation of pre-filter attack surfaces: audio-modality injection, adversarial code comments, and log-embedded payloads all succeed by placing their payloads upstream of any text-level safety defense.

01
mech-interp preprint

From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features

The SAE field has conflated two separable claims: that a decoder vector points in the right direction (geometric recovery) and that the encoder actually fires on relevant inputs (causal activation). This paper runs the first systematic causal audit of SAE evaluation practice and finds the two claims come apart far more often than the field has assumed.

Geometric Recovery vs. Causal Activation Cosine Similarity (geometric recovery) Causal Firing Rate Well-trained SAE (9% inert) Degraded SAE (77% inert) high cosine, never fires
Figure 1: Up to 77% of features passing the cosine ≥ 0.90 recovery threshold in a degraded SAE — and 9% in a well-trained one — never fire when the matched concept is present. Cosine recovery and causal activation are empirically separable claims that existing metrics only measure on the geometric axis.

The pipeline subjects every recovered feature to ablation (zero the feature, measure behavioral change) and steering (activate the feature, measure downstream effect). The 77%/9% causal inertness figures show the problem is severe in degraded SAEs and non-negligible even in well-trained ones. Existing evaluation metrics measure only decoder geometry; they are systematically blind to whether the encoder ever activates. This result should change how SAE papers design evaluations and how we interpret cross-architecture comparisons. Code released at github.com/mohamed-bal/sae-causal-audit.


02
AI security ICSME 2026

The Language of Security: How Prompt Syntax Shapes Code Vulnerability in LLMs

The same security requirement reformulated across five syntactic variants produces up to a 31 percentage-point swing in CWE violation rates — a gap large enough to make prompt syntax the dominant security control in LLM-assisted coding workflows, with model architecture running a distant second.

CWE Violation Rate by Prompt Syntax 0% 20% 40% 60% 58% Decl. 51% Imper. 38% Role 30% CoT 23% Const. 31 pp swing from syntax alone
Figure 2: CWE violation rate across five syntactic prompt formulations, averaged over five frontier LLMs. Declarative instructions produce 58% violation rate; constraint-based formulations reduce this to 23% — a 31 pp swing attributable to syntax alone, not model choice.

Across five major LLMs, 100 coding prompts, and 14 CWE categories, declarative and imperative styles consistently underperform role-based and chain-of-thought instructions by 20–31 pp on high-risk CWEs including CWE-89 (SQL injection) and CWE-79 (XSS). The interaction between model and syntactic variant is significant: a formulation that elicits secure code from GPT-4o may be among the worst choices for Llama 3.3 70B. LLM-assisted coding security audits that test only one prompt formulation systematically underestimate vulnerability rates. Accepted ICSME 2026 main track.


03
mech-interp preprint

ParityTransformer: Interpretability by Design via the Deep Parity Bottleneck

Interpretability in transformers is almost always post-hoc — you train the model, then reverse-engineer what it learned. ParityTransformer inverts the workflow by enforcing an interpretable constraint during training, and the resulting features outperform post-hoc SAEs on every causal metric while matching them on probing.

DPB vs Standard SAE — Causal Metrics Absorption Rate ↓ 28% SAE 12% DPB Steering Accuracy ↑ 61% SAE 74% DPB Causal Stability ↑ 56% SAE 71% DPB
Figure 1: DPB achieves lower feature absorption (12% vs 28%), higher steering accuracy (74% vs 61%), and higher causal intervention stability (71% vs 56%) vs standard SAEs. Probing performance is equivalent — all gains are on causal metrics, which matter for safety applications.

The Deep Parity Bottleneck partitions the residual stream into parity-signed feature subspaces during training, enforcing legibility by construction. Evaluated on GPT-2 and Pythia-1.4B against Anthropic JumpReLU and standard TopK SAEs, DPB models match SAE probing performance but show substantially lower absorption, higher steering accuracy, and more stable causal interventions — with no significant capability degradation. If the result holds at frontier scale, it challenges the assumption that interpretability must be extracted from models post-hoc.


04
mech-interp preprint

Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance

Cross-paper SAE comparisons routinely hinge on autointerpretability scores. This systematic audit shows that pipeline configuration — not SAE architecture — drives the majority of variance in those scores, putting much of the published literature on uncertain empirical footing.

Variance Source: Pipeline vs Architecture Pipeline (corpus, evaluator, template, draw) Architecture Simulation 75% 25% Detection 63% 37% Purity 72% 28% Fuzzing: unreliable across all conditions
Figure 1: Variance decomposition across four autointerpretability metrics on Pythia-160M and Apertus-8B. Pipeline choices (corpus sampling, draw conditions, evaluator model, prompt template) collectively explain more variance than SAE architecture in simulation, detection, and purity; fuzzing is unreliable across all conditions.

Spanning four metrics, two models, and four axes of methodological variation, the paper shows that top-k feature rankings do not stay consistent across corpus and draw conditions — meaning aggregate scores paper over per-feature variance that matters for mechanistic claims. Published SAE architecture comparisons that do not ablate pipeline variables may largely reflect configuration differences rather than genuine architectural signal. Any new paper in this space needs a pipeline ablation to isolate the architectural contribution.


05
mech-interp preprint

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

Features with clear interpretable descriptions often have weak causal effects, while the features that dominate model behavior are often the least well-described. Evaluating SAEs only on activation descriptions misses the half of feature geometry that determines whether steering and ablation will work.

FEGA: Concept vs. Effect Axes are Misaligned Concept Activation Score Functional Effect Score ideal High effect, poor description Good description, weak effect
Figure 1: Feature-Effect Geometry Analysis (FEGA) scatter — SAE features along concept-activation and functional-effect axes. The axes are systematically misaligned: features with high interpretability scores cluster in the bottom-right (good description, weak causal effect), while causally dominant features cluster top-left (high effect, poor description).

FEGA decomposes SAE feature behavior into a concept axis (what activations describe) and an effect axis (how downstream logits shift when the feature fires). Effect geometry predicts steering reliability substantially better than activation descriptions alone. This finding directly complements the causal audit in 2607.12166 and the metric critique in 2607.19386, forming a coherent methodological case against single-axis SAE evaluation. Safety-oriented steering and unlearning applications should select features by effect geometry, not interpretability score.


06
AI security preprint

Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents

Every text-level safety filter acts after audio has been transcribed to text. This attack embeds adversarial instructions as psychoacoustically imperceptible perturbations in the audio waveform itself — placing the payload in the reasoning context before any safety system can act, without replacing the legitimate audio content.

Audio Prompt Injection — Attack Flow Adversarial Audio (≤JND) ASR Layer (Perception) Text context (payload ↑) LLM Agent (executes) Text safety filter ✗ arrives too late — payload already in context
Figure 1: The adversarial payload is embedded in the audio waveform at sub-perceptibility level, transcribed by ASR into the model's reasoning context, and executed before any text-level safety filter can intervene. Text filters arrive after the payload has been promoted to trusted context.

Adversarial instruction strings are embedded as psychoacoustically imperceptible perturbations concurrent with (not replacing) legitimate audio content. The attack exploits the ASR ordering constraint: by the time the model sees text, the payload has already been promoted to trusted context. Evaluated against production-scale multimodal agents, the attack achieves high success rates while remaining below human perceptibility thresholds. Neither users nor acoustic anomaly detectors see anything unusual. Code released.


07
AI security preprint

ToxScreen: Detecting Whether an LLM Has Been Poisoned

Backdoor defense research has lacked a benchmark that matches the realistic white-box threat model — defender has weight access but no training data, no trusted reference model, and no prior on trigger type. ToxScreen is the first systematic substrate for evaluating trigger-recovery defenses under these constraints.

ToxScreen — ~800 Backdoored Models Attack Objectives Trigger Mechanisms Poison Rate 0.1%–5% Model Scale 125M–7B Training Method FT/LoRA/RLHF White-box Defender Threat Model Weight access · No training data · No reference model · No trigger prior
Figure 1: ToxScreen benchmark — ~800 backdoored models varying across five axes evaluated under a realistic white-box threat model where the defender has weight access but no training data, no trusted reference model, and no prior on trigger type.

The benchmark spans ~800 backdoored models covering attack objectives, trigger mechanisms (word/phrase/style triggers), poisoning rates 0.1%–5%, model scales 125M–7B, and training mechanisms including full fine-tuning, LoRA, and RLHF-based backdoor insertion. This is the first systematic substrate for evaluating trigger-recovery defenses under a realistic threat model. Competitive baseline results are expected to follow shortly.


08
mech-interp preprint

Where Steering Signals Come From: Activation Source Selection in Activation Steering

Source position — where in the context you extract the steering vector — has been treated as an unexamined hyperparameter. This paper shows it is the decisive variable: execution-boundary states substantially outperform all alternatives across models and tasks.

Steering Success by Source Position Source Position 31% Random 38% Last Token 42% Post-action 71% Exec. Boundary 50%
Figure 1: Steering vector effectiveness at different source positions — averaged across three instruction-tuned models and four task families. Execution-boundary positions (just before the model produces the target behavior) consistently outperform random positions, last tokens, and post-action states.

Holding the downstream intervention fixed and varying only the source position, execution-boundary states — positions just before the model produces the target behavior — yield substantially stronger signals than random positions, last tokens, or post-action states. The result holds across three instruction-tuned models and four task families, turning source selection from a poorly understood hyperparameter into a principled design choice with direct implications for safety steering and unlearning applications.

Items 9 – 14 · Also notable
09
AI security preprint

The Illusion of Secure LLM Code: How Iterative Reprompting Undermines Security

CWE Violations Over Reprompting Iterations Reprompting Iteration Cumulative CWEs 0 1 2 3 4 5 +37.6%
Figure 4: Cumulative critical CWE violations increase +37.6% after 5 reprompting iterations as functional objectives supersede initial security constraints.

Most LLMs produce low-risk code on initial request, but iterative agentic reprompting causes critical CWE violations to increase by +37.6% after 5 iterations as the model increasingly prioritizes the most recently specified functional objectives over earlier security constraints. The pattern is consistent across GPT-4o, Claude 3.7 Sonnet, and three other frontier models.


10
AI security ICLR 2026 Workshop

Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety

MAP-Elites — Per-Model Safety Geometry Plain ROT13 Leet Base64 Direct Hypothetical Roleplay GPT ASR 0.8 GPT ASR 0.7 C: 0.4 C: 0.4 C: 0.4
Figure 5: MAP-Elites behavioral grid — each cell is a strategy × encoding combination; red = high ASR. GPT-4o-mini and Gemini show concentrated vulnerabilities; Claude shows uniformly low ASR (max 0.4) across all cells.

Applies MAP-Elites evolutionary search to LLM red-teaming over interpretable semantic attack strategies. The approach surfaces per-model safety geometry — which strategy-encoding pairs are productive and which are dead ends — enabling targeted hardening. ICLR 2026 Workshop on Agents in the Wild.


11
AI security preprint

LogInject: Passive Prompt Injection via Security Operations Log Context

LogInject Attack Flow HTTP UA / DNS / SMTP Server Log Entry SOC Agent LLM Context Attack Executes Text filter ✗ bypassed 12,847-entry LogInject-1.0 benchmark · all 4 attack objectives succeed No user interaction required · structural metadata bypasses free-text filters
Figure 7: Adversarial payloads in HTTP User-Agent strings, DNS TXT records, and SMTP headers propagate into server logs; SOC agents treat log content as authoritative, executing all four attack objectives with no user interaction.

Adversarial payloads embedded in standard log metadata (HTTP User-Agent, DNS TXT, mail headers) reach SOC agent LLMs as trusted operational context, succeeding at all four attack objectives across 12,847 log entries. Most IPI defenses filter free-text fields and miss payloads in structural metadata — opening a blind spot in AI-assisted SOC pipelines.


12
mech-interp preprint

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Bias as Geometric Subspace in Judge Activations Clean inputs (on manifold) position bias verbosity bias authority bias Steerable in both directions
Figure 1: Biased inputs occupy type-specific low-dimensional subspaces that sharpen with model depth; steering along the bias subspace drives scoring in either direction across seven judges and seven bias types.

Covers seven judges, seven bias types (position, verbosity, authority, format, etc.), and nine benchmarks. Biases are geometrically structured, type-specific, and causally controllable via hidden-state steering — establishing judge alignment as a mechanistic problem amenable to intervention rather than just an instruction-tuning shortfall.


13
AI security preprint

ALIBI: Adaptive Agentic Attacks on LLM-Based Vulnerability Detectors via Adversarial Code Comments

ALIBI — Adaptive Comment Injection Loop Vulnerable Code + Adv. Comments LLM Detector (chain-of-thought) Evasion achieved (detector fooled) refine comments ← feedback
Figure 1: ALIBI loop — agent iteratively inserts and refines adversarial natural-language comments against four LLM-based vulnerability detectors, exploiting the gap between comment semantics and code security semantics.

A coding agent iteratively inserts adversarial natural-language comments in real-world vulnerable code from vulnerability-fixing commits, exploiting that comments are parsed as natural language by LLM detector chain-of-thought but carry no code security semantics. Evaluated against four detectors from specialized reasoning models to frontier multi-agent systems; source-code comments emerge as a previously understudied attack surface.


14
mech-interp preprint

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders

Inner-Product vs. Cosine SAE Feature Use Inner-Product SAE Content features Norm detectors ~20–30% wasted slots Cosine SAE All slots → content features norm dependence ≈ 0 after training
Figure 8: Inner-product SAEs allocate a fraction of dictionary slots to norm detectors (a quantity transformers discard via sublayer normalization); cosine training converges to near-zero magnitude dependence, filling all slots with content-aligned features.

Standard SAE inner-product scores depend on input norm — a quantity sublayer normalization discards from the residual stream, causing wasted dictionary slots. Replacing inner-product with a learned cosine blend always converges to near-zero magnitude dependence, reclaiming those slots for content-aligned features with measurable semantic alignment improvement at matched reconstruction loss.

Watchlist
← all Research Radar issues · gussand · source