📡 Research Radar

Daily Report — 24 September 2026

Window: Sept 22–24, 2026
Sources: OpenReview · EMNLP 2026 · USENIX Security 26 · ICLR 2026 · arXiv cs.CL/cs.LG/cs.CR/cs.AI · LessWrong / Alignment Forum
3 peer-reviewed 7 preprints 0 forum/blog

ITEMS 4 – 10 · MID-TIER

06 / AI SECURITY / USENIX SECURITY 2026

prompt-injection on-policy-distillation USENIX Security 26

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

(Research team) — 35th USENIX Security Symposium 2026 (peer-reviewed)

Step 1: On-Policy Sampling Current defender generates completions for injection prompts → token-level reward signal Step 2: Distillation Update Teacher provides token-level feedback; close the adaptive loop no manual attack enumeration adaptive

SecOPD two-step loop: on-policy sampling generates completions under current injection pressure; token-level distillation closes the adaptive loop without manual attack enumeration. Outperforms static fine-tuning defences in utility-under-attack metrics.

SecOPD frames prompt injection defence as on-policy distillation: the current defender model generates completions for a live distribution of injection prompts, and a teacher provides token-level reward signals, closing the adaptive loop without requiring manual attack corpus construction. This approach sharply reduces attack success rates against adaptive adversaries who iteratively update payloads to defeat defences, and outperforms static fine-tuning baselines in utility-under-attack metrics across multiple task settings; published at USENIX Security 26, the strongest security venue represented this cycle.

09 / APPLIED MECH INTERP / PREPRINT

steering-vectors SAE applied-interp

Disentangling Steering Vectors

Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi, Hisashi Kashima — arXiv preprint, September 2026

DiffMeans composite SAE concept A e.g. sentiment concept B e.g. topic concept C e.g. style Targeted single-concept steering

Composite DiffMeans steering vector decomposed via Steering Vector Dissection (SAE on per-token activation deltas) into semantically distinct single-concept components. Targeted interventions using individual components reduce side-effect co-activation versus the composite direction.

8 entries removed on 2026-09-29 as repeats of earlier reports: 2508.06601 (first covered 2026-09-02), 2609.20412 (first covered 2026-09-18), 2607.26654 (first covered 2026-09-01), 2609.03887 (first covered 2026-09-08), 2609.15886 (first covered 2026-09-17), 2608.27504 (first covered 2026-09-01), 2609.04022 (first covered 2026-09-08), 2609.04721 (first covered 2026-09-08).

Difference-in-means steering vectors conflate multiple semantic and stylistic concepts, causing unpredictable side effects; Steering Vector Dissection trains a Sparse Autoencoder on per-token steering vectors derived from paired activation datasets, then uses decoder columns as candidate concept directions, filtered by interpretability and steering experiments. The framework decomposes composite DiffMeans vectors into semantically clean single-concept components that transfer more reliably across prompts and reduce unwanted concept co-activation — a practical applied-interp tool for controlled, auditable model steering.

Notes

Items 2 and 5 form a productive tension: both stress-test midtraining interventions and find them brittle in different ways — Geodesic's result shows even large midtraining investments collapse under minimal competing fine-tuning; the inoculation paper finds neologism-based context segregation leaks under distribution shift. Item 3 (Constitutional Midtraining) shows partial rebuttal: content-presence effects on specific behaviours survive benign fine-tuning. Together they define the current boundary of what midtraining can and cannot guarantee. Items 4 and 7 advance the mechanistic safety case: training method determines circuit structure (EMNLP 2026); those circuits can be located and ablated to reduce jailbreak ASR by 80% (preprint).

← all Research Radar issues · gussand · source