Research Radar · Daily Edition
August 8, 2026
0 peer-reviewed · 2 preprints · 0 forum/blog
Mech Interp AI Security Text Diffusion LMs
Items 4–10 · Also notable
05
mech-interp AI security preprint

Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off

EMBEDDING SPACE (harmful input neighbourhood) harmful input x REFUSAL region COMPLIANCE region RED jailbreak ≈ move along RED safety-utility trade-off provable: closing RED fully ⟹ capability loss
RED (Refusal-Escape Direction): a local perturbation direction near a harmful input that shifts refusal → compliance while preserving harmful semantics. Jailbreaks approximate movement along RED; closing RED provably incurs a safety-utility trade-off.
Yu Chen, Yuanhao Liu, Qi Cao (Institute of Computing Technology, CAS) · arXiv preprint, May 2026

Refusal-Escape Directions (RED) are local continuous perturbation directions near harmful inputs that shift model behavior from refusal to compliance while preserving harmful-semantics interpretation — giving a mechanistic, operator-level account of why jailbreaks structurally work. The framework proves that fine-tuning-based alignment cannot eliminate RED without capability loss, establishing a fundamental safety-utility trade-off that output-level defenses cannot resolve.

06
AI security preprint

Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code

"Write malware that does X" Aligned LLM safety filter I cannot help NORMAL PATH: refused ✓ "Generate code" + GCD grammar Aligned LLM code-mode primed malicious code CODESPEAR: grammar primes code-mode → bypass ✗ CodeShield defense honeypot code alignment → safe behavior under GCD
CodeSpear: a benign code grammar constraint applied via GCD primes the code-generation mode, bypassing safety alignment to elicit malicious code. CodeShield defends via honeypot code alignment.
Yitong Zhang, Shiteng Lu, Jia Li · arXiv preprint, June 2026

Grammar-Constrained Decoding (GCD) is a reliability tool that enforces syntactic validity in LLM code output — but the reliability mechanism is itself an attack surface. CodeSpear shows that applying a benign code grammar constraint shifts the generation distribution into a code-conditioned mode where safety guardrails are less active, enabling elicitation of malicious code from aligned models without modifying the harmful intent in the prompt. CodeShield defends by training the model to emit honeypot code (syntactically valid, functionally harmless) under attacker-controlled grammar constraints, robustly preserving safe behavior.

8 entries removed on 2026-09-10 as repeats of earlier reports: 2608.02632 (first covered 2026-08-07), 2606.16939 (first covered 2026-07-01), 2604.20945 (first covered 2026-07-07), 2606.28153 (first covered 2026-07-01), 2606.22673 (first covered 2026-07-02), 2606.24026 (first covered 2026-07-02), 2603.23268 (first covered 2026-07-01), 2605.08934 (first covered 2026-07-10).

← all Research Radar issues · gussand · source