Window: Aug 24–26, 2026 + uncovered peer-reviewed work
0 peer-reviewed · 3 preprints · 0 forum/blog
Mech Interp·AI Security·Text Diffusion LMs
Items 4–10 · Also notable
Figure 1 · FAR.AI AI Security Leaderboard: cost to find a universal jailbreak against CBRNE / offensive cybersecurity requests. Grok 4.5 ($58) and Gemini 3.1 Pro ($278) are broken; Claude Fable 5 and GPT-5.6 Sol exceed $14,200 and were not broken. Over 100× gap.
FAR.AI defines the Minimal Standard for Safeguards v1.0 across CBRNE and offensive cybersecurity and finds >100× gap between frontier models: Grok 4.5 yields a universal jailbreak for ~$58, Gemini 3.1 Pro for ~$278, while Claude Fable 5 and GPT-5.6 Sol resist all tested jailbreak families at >$14,200 search cost. Models strong on narrow benchmarks are shown to fail the Standard in adjacent hazard domains.
Sequential activation patching: patches applied at each CoT token position, aggregated via PoS-guided analysis to locate which positions carry peak causal influence on the final answer.
Murat Dura, Serkan Öztürk, Selma Tekir. Standard activation patching patches a single static position — insufficient for CoT which unfolds over many tokens. Sequential activation patching applies patches at each CoT token and uses Part-of-Speech-guided aggregation to identify attention heads that carry causal signals propagating to final-answer computation.
Mechanistic Tomography: the forward pass under interventions modeled as y = Ax + w; A encodes the intervention design, x is the mechanism map to recover, w absorbs noise and nonlinear residuals. Identifiability conditions derived analytically.
Vijay Erramilli (24 pp., 13 figures). Formulates mechanistic interpretability as a designed measurement inverse problem y = Ax + w, where A is the intervention matrix and x is the mechanism map to recover. Derives identifiability conditions from activation-patching designs, unifying control theory and mechanistic interpretability in a single formal framework.
Notes
Reports gap: Last radar was August 7. This run covers August 8–26 with emphasis on August 24–26 fresh submissions plus ICLR 2026 and ACL 2026 peer-reviewed work not previously included.
Pattern flag: Items 01, 02, and 10 each attack diffusion LLM safety from a distinct angle — mechanistic neuron pruning (01), architectural bidirectionality exploitation (02), and black-box schema learning (10). Three concurrent independent attack vectors suggest the dLLM safety research community is entering an adversarial phase analogous to autoregressive jailbreak research circa 2023.