Figure 1: Per-layer transcoder (PLT) attribution graph on Qwen3-4B for a deceptive completion. Deception-specific features (red, larger radius ∝ weight) exert 2–3× stronger causal influence on the output than honest features (blue, dashed connections), and direct steering of these features flips the response between deceptive and honest in a predictable, controlled manner.
What does deception look like inside a language model? Using per-layer transcoders — the same attribution-graph framework Anthropic employs for internal circuit analysis — this paper builds the first mechanistic map of LLM deception, finding a concentrated dictionary of deception-specific features that dominate the causal graph when the model is lying.
Per-layer transcoders (PLTs) applied to Qwen3-4B construct attribution graphs capturing feature activations and inter-feature dependencies across layers. Comparing honest and deceptive completions on the same underlying facts, the paper identifies deception-related features that (1) activate significantly more strongly in the deceptive regime, (2) exert 2–3× greater causal influence on the output distribution than honest-content features, and (3) when directly steered via activation addition, produce predictable, controlled shifts between deceptive and non-deceptive responses. Honest-content features are correspondingly suppressed in the deceptive completion graph. This mechanistic handle on deception is independent of access to training data or reward signals, and the identified features provide interpretable, targetable circuit components for deception detection and mitigation.
Figure 1: RARE separates expert routing (which expert activates, teal) from representation steering (how the hidden state is transformed, violet) — enabling independent interpretability and targeted intervention on each component in MoE language models.
Accepted to the Actionable Interpretability Workshop at COLM 2026. Standard MoE forward passes couple expert routing with representation transformation, making it difficult to interpret or steer either in isolation. RARE disentangles the two operations — providing separate handles for "which expert" vs. "how the representation is steered" — enabling cleaner interpretability experiments and targeted interventions in MoE architectures.
Notes
Only 4 genuinely relevant uncovered items found; no padding. The Aug 23–25 window had sparse new submissions — consistent with the late-August lull noted in the Aug 20 and Aug 23 radars. All four items were missed by prior runs.
Item #1 (BSF, Goodfire) is the methodological standout of this window. The 1D→subspace shift directly challenges the SAE paradigm underpinning most mech-interp work of the past two years. If the results hold broadly, prior conclusions attributed to specific "features" may need revisiting through the manifold lens. Flag for weekly roundup.
Items #2 and #3 form a jailbreak–deception mechanistic pair. Both apply attribution-graph methodology (the same tool Anthropic uses internally) to reveal that safety failures — whether jailbreaks (#2) or deliberate deception (#3) — have identifiable, mechanistically legible substrates: suppressed safety nodes, emergent attack/deception features, and predictable causal pathways. Flag for weekly roundup.
Item #4 (RARE) is the only peer-reviewed entry this window (COLM 2026 Actionable Interpretability Workshop). Limited public details available from external sources; the workshop status confirms full peer review.
No text diffusion LM papers this window.
2 entries removed on 2026-09-10 as repeats of earlier reports: 2606.25234 (first covered 2026-08-24), 2607.07903 (first covered 2026-07-12).