Research Radar

Daily Digest — August 18, 2026

Daily · August 8–18, 2026 · sources: TrustNLP @ ACL 2026 · arXiv cs.CL/cs.LG/cs.CR/cs.AI
1 peer-reviewed · 3 preprints · 0 forum/blog
Mech Interp AI Security Text Diffusion LMs
Items 4 – 10 · Also notable
05 · mech-interp · AI security · preprint
mech-interp AI security preprint

Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning

Zirui Song et al. · arXiv, August 2026 · 398 public unlearned models audited
J-ACCESS AUDIT — 398 PUBLIC UNLEARNED MODELS Pre-attack J-Access accessibility score Recovery speed (fine-tuning steps) gold 0 low mid high most models retain access above gold baseline
Figure 1 · J-Access audit of 398 public unlearned models — pre-attack accessibility (Jacobian lens) predicts recovery speed; most models retain internal access above the retain-only gold baseline despite passing surface-behavior unlearning tests.

J-Access uses the Jacobian lens to map intermediate representations into vocabulary space and measure how often target concepts remain accessible along the output pathway — without triggering surface refusal. Auditing 398 public unlearned models spanning eight unlearning methods, the study finds most retain access above the retain-only gold level. Pre-attack accessibility predicts recovery speed and extent at the model level, establishing it as a diagnostic metric for whether fine-tuning will restore "unlearned" knowledge.

06 · mech-interp · AI security · TrustNLP @ ACL 2026 (peer-reviewed)
TrustNLP @ ACL 2026 mech-interp AI security

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs

Shivam Ratnakar, Kartikeya Vats · TrustNLP 2026 @ ACL 2026 (peer-reviewed)
CLS — REFUSAL DIRECTION GEOMETRY IN ACTIVATION SPACE Benign queries Malicious queries refusal direction (CLS) safe vs unrestricted system prompt contrast CLS project refusal direction onto vocabulary → upweights refusal tokens no activation patching
Figure 1 · CLS isolates the refusal direction by contrasting hidden states from safe vs. unrestricted system prompts; malicious and benign queries form separable linear clusters — safety is a manipulable linear feature, not a deep semantic decision.

Contrastive Logit Steering (CLS) isolates the "refusal direction" by contrasting hidden states from safe and unrestricted system prompts, then projects this direction onto the vocabulary to upweight refusal tokens — operating entirely at the output-distribution level without modifying activations. Demonstrates that malicious and benign queries form distinct, linearly separable clusters in activation space, that safety compliance is a manipulable linear feature, and that removing it is geometrically trivial. Peer-reviewed at TrustNLP @ ACL 2026.

09 · text diffusion · preprint
text diffusion preprint

Commitment Before Realization: When Classifier-Free Guidance Becomes Unnecessary in Masked Diffusion Language Models

Fan Zhou, Weitian Wang, Tim Van de Cruys · arXiv, August 8, 2026
COMMITMENT HORIZON — CFG VALUE OVER DECODING STEPS Decoding step (masked token index) CFG value commitment horizon (prompt A) commitment horizon (prompt B) no CFG benefit (prompt C) high mid none
Figure 1 · Commitment horizon curves per prompt — prompt A commits early (CFG can be dropped mid-decoding), prompt B benefits through the middle, prompt C sees no CFG benefit throughout; guidance need is highly prompt-specific.

Defines the "commitment horizon" — the earliest decoding step from which switching to base-model (no CFG) reduces final constraint-satisfaction success by no more than a chosen tolerance. Guidance dependence is highly prompt-specific: many prompts succeed without CFG, others see no benefit or are harmed by it, and for those that do benefit, the gain concentrates early. Result: dynamic CFG scheduling can match full-CFG quality at substantially reduced compute for masked dLLM inference.

10 · mech-interp · preprint
mech-interp preprint

Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models

Hao Ai (Tsinghua University) · arXiv, August 2026
MEAN-FIELD CoT — CLUE DISCOVERY FRACTION vs. REASONING STEP CoT reasoning step Fraction of clues discovered 0 0.5 1.0 ODE (mean-field) ● empirical avg emergence threshold
Figure 1 · Mean-field ODE for clue-discovery fraction over CoT steps — theoretical curve fits empirical averages; the framework predicts emergence thresholds and convergence without simplifying model architecture.

Formulates Chain-of-Thought reasoning as a guided discovery process on a latent clue graph and derives a 1D ODE for the fraction of clues discovered per step using the mean-field approximation. Clue tokens are identified via normalized surprisal of a student LLM on teacher outputs; statistical regularities are averaged over many chains. First mean-field theoretical framework for CoT dynamics that predicts emergence thresholds and convergence properties without simplifying model architecture or using physical system analogies.

6 entries removed on 2026-09-10 as repeats of earlier reports: 2608.09867 (first covered 2026-08-13), 2608.07430 (first covered 2026-08-11), 2608.10172 (first covered 2026-08-13), 2608.08168 (first covered 2026-08-13), 2608.10530 (first covered 2026-08-13), 2608.04018 (first covered 2026-08-11).

← all Research Radar issues · gussand · source