Zirui Song et al. · arXiv, August 2026 · 398 public unlearned models audited
Figure 1 · J-Access audit of 398 public unlearned models — pre-attack accessibility (Jacobian lens) predicts recovery speed; most models retain internal access above the retain-only gold baseline despite passing surface-behavior unlearning tests.
J-Access uses the Jacobian lens to map intermediate representations into vocabulary space and measure how often target concepts remain accessible along the output pathway — without triggering surface refusal. Auditing 398 public unlearned models spanning eight unlearning methods, the study finds most retain access above the retain-only gold level. Pre-attack accessibility predicts recovery speed and extent at the model level, establishing it as a diagnostic metric for whether fine-tuning will restore "unlearned" knowledge.
Figure 1 · CLS isolates the refusal direction by contrasting hidden states from safe vs. unrestricted system prompts; malicious and benign queries form separable linear clusters — safety is a manipulable linear feature, not a deep semantic decision.
Contrastive Logit Steering (CLS) isolates the "refusal direction" by contrasting hidden states from safe and unrestricted system prompts, then projects this direction onto the vocabulary to upweight refusal tokens — operating entirely at the output-distribution level without modifying activations. Demonstrates that malicious and benign queries form distinct, linearly separable clusters in activation space, that safety compliance is a manipulable linear feature, and that removing it is geometrically trivial. Peer-reviewed at TrustNLP @ ACL 2026.
Fan Zhou, Weitian Wang, Tim Van de Cruys · arXiv, August 8, 2026
Figure 1 · Commitment horizon curves per prompt — prompt A commits early (CFG can be dropped mid-decoding), prompt B benefits through the middle, prompt C sees no CFG benefit throughout; guidance need is highly prompt-specific.
Defines the "commitment horizon" — the earliest decoding step from which switching to base-model (no CFG) reduces final constraint-satisfaction success by no more than a chosen tolerance. Guidance dependence is highly prompt-specific: many prompts succeed without CFG, others see no benefit or are harmed by it, and for those that do benefit, the gain concentrates early. Result: dynamic CFG scheduling can match full-CFG quality at substantially reduced compute for masked dLLM inference.
Figure 1 · Mean-field ODE for clue-discovery fraction over CoT steps — theoretical curve fits empirical averages; the framework predicts emergence thresholds and convergence without simplifying model architecture.
Formulates Chain-of-Thought reasoning as a guided discovery process on a latent clue graph and derives a 1D ODE for the fraction of clues discovered per step using the mean-field approximation. Clue tokens are identified via normalized surprisal of a student LLM on teacher outputs; statistical regularities are averaged over many chains. First mean-field theoretical framework for CoT dynamics that predicts emergence thresholds and convergence properties without simplifying model architecture or using physical system analogies.