06 / AI SECURITY / USENIX SECURITY 2026
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
(Research team) — 35th USENIX Security Symposium 2026 (peer-reviewed)
SecOPD two-step loop: on-policy sampling generates completions under current injection pressure; token-level distillation closes the adaptive loop without manual attack enumeration. Outperforms static fine-tuning defences in utility-under-attack metrics.
SecOPD frames prompt injection defence as on-policy distillation: the current defender model generates completions for a live distribution of injection prompts, and a teacher provides token-level reward signals, closing the adaptive loop without requiring manual attack corpus construction. This approach sharply reduces attack success rates against adaptive adversaries who iteratively update payloads to defeat defences, and outperforms static fine-tuning baselines in utility-under-attack metrics across multiple task settings; published at USENIX Security 26, the strongest security venue represented this cycle.