Research Radar

Daily · August 26, 2026

Window: Aug 24–26, 2026 + uncovered peer-reviewed work
0 peer-reviewed · 3 preprints · 0 forum/blog
Mech Interp · AI Security · Text Diffusion LMs
Items 4–10 · Also notable
FAR.AI Leaderboard: Cost to Find a Universal Jailbreak (CBRNE + Cybersec) Grok 4.5 Gemini 3.1 Pro GPT-5.6 Sol Claude Fable 5 $58 $278 >$14,200 — not broken → >$14,200 — not broken → FAR.AI Minimal Standard for Safeguards v1.0 · >100× gap in robustness
Figure 1 · FAR.AI AI Security Leaderboard: cost to find a universal jailbreak against CBRNE / offensive cybersecurity requests. Grok 4.5 ($58) and Gemini 3.1 Pro ($278) are broken; Claude Fable 5 and GPT-5.6 Sol exceed $14,200 and were not broken. Over 100× gap.
05
AI security preprint · Aug 5, 2026

AI Security Leaderboard: Methodology, Results and Minimal Standard

FAR.AI defines the Minimal Standard for Safeguards v1.0 across CBRNE and offensive cybersecurity and finds >100× gap between frontier models: Grok 4.5 yields a universal jailbreak for ~$58, Gemini 3.1 Pro for ~$278, while Claude Fable 5 and GPT-5.6 Sol resist all tested jailbreak families at >$14,200 search cost. Models strong on narrow benchmarks are shown to fail the Standard in adjacent hazard domains.

Sequential Activation Patching across Chain-of-Thought Token Positions Prompt Q: ... Chain-of-Thought reasoning tokens (patched sequentially) tok₁ tok₂ tok₃ ··· tokₙ VERB NOUN NUM NOUN ← PoS-guided aggregation of patch effects → Final answer Peak causal influence here
Sequential activation patching: patches applied at each CoT token position, aggregated via PoS-guided analysis to locate which positions carry peak causal influence on the final answer.
06
mech interp preprint · Aug 23, 2026

Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

Murat Dura, Serkan Öztürk, Selma Tekir. Standard activation patching patches a single static position — insufficient for CoT which unfolds over many tokens. Sequential activation patching applies patches at each CoT token and uses Part-of-Speech-guided aggregation to identify attention heads that carry causal signals propagating to final-answer computation.

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability y = Ax + w measurements intervention design matrix mechanism map to recover noise + nonlinear res. Identifiability Derived analytically from activation patching designs Bridges control theory → mechanistic interpretability · 24 pp. · 13 figures
Mechanistic Tomography: the forward pass under interventions modeled as y = Ax + w; A encodes the intervention design, x is the mechanism map to recover, w absorbs noise and nonlinear residuals. Identifiability conditions derived analytically.
07
mech interp preprint · Aug 2026

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

Vijay Erramilli (24 pp., 13 figures). Formulates mechanistic interpretability as a designed measurement inverse problem y = Ax + w, where A is the intervention matrix and x is the mechanism map to recover. Derives identifiability conditions from activation-patching designs, unifying control theory and mechanistic interpretability in a single formal framework.

Notes

Reports gap: Last radar was August 7. This run covers August 8–26 with emphasis on August 24–26 fresh submissions plus ICLR 2026 and ACL 2026 peer-reviewed work not previously included.

Pattern flag: Items 01, 02, and 10 each attack diffusion LLM safety from a distinct angle — mechanistic neuron pruning (01), architectural bidirectionality exploitation (02), and black-box schema learning (10). Three concurrent independent attack vectors suggest the dLLM safety research community is entering an adversarial phase analogous to autoregressive jailbreak research circa 2023.

7 entries removed on 2026-09-10 as repeats of earlier reports: 2608.07430 (first covered 2026-08-11), 2507.11097 (first covered 2026-07-07), 2608.10172 (first covered 2026-08-13), 2508.13650 (first covered 2026-07-19), 2606.18383 (first covered 2026-08-04), 2606.06333 (first covered 2026-07-02), 2606.04027 (first covered 2026-07-02).

← all Research Radar issues · gussand · source