📡 Research Radar
Daily · September 23, 2026
2 peer-reviewed · 3 preprints · 0 forum/blog
Pretraining Safety · AI Security · Mech Interp
Items 4–10 · Also notable

05
AI security jailbreak-defense preprint

Dynamic Deep Prompt Optimization for Defending Against Jailbreak Attacks on LLMs

DDPO: Dynamic Deep Prompt Optimization Jailbreak Prompt LLM Layers (feature ext.) hidden states Lightweight MLP Dynamic Defense Embedding per-query adaptive · no additional fine-tuning required
Figure 5: DDPO pipeline — LLM's own intermediate hidden states → lightweight MLP → dynamic per-query defensive embedding; adapts defenses to each input rather than relying on a static prefix.

DDPO (Dynamic Deep Prompt Optimization) is the first jailbreak defense based on deep prompt optimization: the target LLM's own intermediate layers serve as feature extractors, feeding a lightweight MLP that dynamically generates defensive embeddings per query without expensive fine-tuning — outperforming static prompt-based defenses across multiple attack types.


06
AI security encrypted-inference preprint

HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference

HE-Guardrail: Jailbreak Defense Over Encrypted Inference CLIENT prompt (plaintext) HE encrypt ciphertext SERVER (no plaintext) HE Guardrail runs over ciphertext BLOCK jailbreak / PASS safe → encrypted LLM inference Safe encrypted LLM response (or blocked)
Figure 6: HE-Guardrail: the server evaluates guardrail logic entirely over homomorphically encrypted ciphertexts — the first framework enabling jailbreak detection without ever decrypting the client's prompt.

In HE-LLM inference the server cannot inspect plaintext prompts — but this also shields adversarial jailbreak payloads from detection. HE-Guardrail is the first framework that evaluates guardrail mechanisms entirely over homomorphically encrypted data, closing a novel attack surface unique to privacy-preserving LLM serving.


07
ICML 2026 MI Workshop mech-interp causal-circuits

Topographic Training Concentrates Causal Circuits Without Improving Neuron Monosemanticity

Topographic vs Standard Training Causal Sufficiency Std Topo 2.79× Neuron Monosemanticity Std Topo ≈ same SAE L0 sparsity ↓11% · dead-feature fraction ↑19× · circuit-level effect, not neuron-level
Figure 7: Topographic clusters are 2.79× more causally sufficient than random unit sets; SAE L0 sparsity decreases 11% and dead-feature fraction rises 19-fold, yet neuron monosemanticity scores are unchanged — topographic pressure acts at circuit level.

Applying a spatial-locality TopoLoss makes topographic clusters 2.79× more causally sufficient than random unit sets while SAE L0 sparsity decreases 11% and dead-feature fraction rises 19-fold, yet standard neuron monosemanticity scores remain unchanged — separating circuit-level causal concentration from neuron-level feature disentanglement as distinct interpretability properties.


08
mech-interp formal-verification preprint

Certified Mechanistic Interpretability: Lifting Single-Input Findings to Bounded Neighbourhoods

CPZ Propagation: Single Input → Neighbourhood Certificate single input x₀ "head H → subj S" CPZ x₀ ε ε-neighbourhood Certificate "H→S holds for all x ∈ N(x₀, ε)"
Figure 8: CPZ propagation lifts a single-input finding (e.g. "attention head H attends to subject S") to a certified ε-neighbourhood; the recursive Jacobian zonotope extends certificates across layer depth without per-layer generator growth.

Constrained polynomial-zonotope (CPZ) propagation through transformer blocks lifts mechanistic-interpretability observations from a single input to certified statements over a bounded neighbourhood, converting brittle single-point findings into robust mechanistic claims; three internal-attention queries (top-k stability, evidence mass, attention entropy) are formulated as tractable programs over the softmax simplex.


09
mech-interp SAE AI-agents preprint

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

Agent vs. Expert Performance — SAEScientist-Bench Feature Sep. Causal Steering (high) (weak) Expert Agent
Figure 9: Frontier agents approach expert levels on contrastive feature separation within a 131K-feature Gemma Scope dictionary but lag substantially on causal-generation steering across 10 agent configurations and 20 tasks.

SAEScientist-Bench evaluates frontier AI agents on autonomous SAE interpretability research: given a target concept, agents design contrastive probes and navigate a 131K+ feature Gemma Scope dictionary to discover optimal features. Agents demonstrate genuine discovery capabilities — approaching expert levels on contrastive feature separation — but lag substantially on causal-generation steering and frequently misinterpret experimental measurements.


← all Research Radar issues · gussand · source