Figure 5: DDPO pipeline — LLM's own intermediate hidden states → lightweight MLP → dynamic per-query defensive embedding; adapts defenses to each input rather than relying on a static prefix.
DDPO (Dynamic Deep Prompt Optimization) is the first jailbreak defense based on deep prompt optimization: the target LLM's own intermediate layers serve as feature extractors, feeding a lightweight MLP that dynamically generates defensive embeddings per query without expensive fine-tuning — outperforming static prompt-based defenses across multiple attack types.
Figure 6: HE-Guardrail: the server evaluates guardrail logic entirely over homomorphically encrypted ciphertexts — the first framework enabling jailbreak detection without ever decrypting the client's prompt.
In HE-LLM inference the server cannot inspect plaintext prompts — but this also shields adversarial jailbreak payloads from detection. HE-Guardrail is the first framework that evaluates guardrail mechanisms entirely over homomorphically encrypted data, closing a novel attack surface unique to privacy-preserving LLM serving.
Figure 7: Topographic clusters are 2.79× more causally sufficient than random unit sets; SAE L0 sparsity decreases 11% and dead-feature fraction rises 19-fold, yet neuron monosemanticity scores are unchanged — topographic pressure acts at circuit level.
Applying a spatial-locality TopoLoss makes topographic clusters 2.79× more causally sufficient than random unit sets while SAE L0 sparsity decreases 11% and dead-feature fraction rises 19-fold, yet standard neuron monosemanticity scores remain unchanged — separating circuit-level causal concentration from neuron-level feature disentanglement as distinct interpretability properties.
Figure 8: CPZ propagation lifts a single-input finding (e.g. "attention head H attends to subject S") to a certified ε-neighbourhood; the recursive Jacobian zonotope extends certificates across layer depth without per-layer generator growth.
Constrained polynomial-zonotope (CPZ) propagation through transformer blocks lifts mechanistic-interpretability observations from a single input to certified statements over a bounded neighbourhood, converting brittle single-point findings into robust mechanistic claims; three internal-attention queries (top-k stability, evidence mass, attention entropy) are formulated as tractable programs over the softmax simplex.
Figure 9: Frontier agents approach expert levels on contrastive feature separation within a 131K-feature Gemma Scope dictionary but lag substantially on causal-generation steering across 10 agent configurations and 20 tasks.
SAEScientist-Bench evaluates frontier AI agents on autonomous SAE interpretability research: given a target concept, agents design contrastive probes and navigate a 131K+ feature Gemma Scope dictionary to discover optimal features. Agents demonstrate genuine discovery capabilities — approaching expert levels on contrastive feature separation — but lag substantially on causal-generation steering and frequently misinterpret experimental measurements.