Research Radar
Daily · August 25, 2026
1 workshop accepted · 3 preprints · 0 forum/blog
Mech Interp · AI Security · Text Diffusion LMs
03
mech-interp AI security preprint

Transcoders for Investigating Deception in Language Models

Qwen3-4B — Per-Layer Transcoder (PLT) Attribution Layer 8 PLT f₁ f₂ f₃ Layer 16 PLT f₄ f₅ f₆ Layer 24 PLT f₇ f₈ f₉ Attribution Graph — deceptive completion truth w=0.3 fact w=0.2 deceive w=0.82★ conceal w=0.61★ Deceptive output Key finding • Deception features activate 2–3× more strongly • Honest features suppressed in deceptive regime • Steering deception features flips output predictably (same transcoder tools Anthropic uses for Claude)
Figure 1: Per-layer transcoder (PLT) attribution graph on Qwen3-4B for a deceptive completion. Deception-specific features (red, larger radius ∝ weight) exert 2–3× stronger causal influence on the output than honest features (blue, dashed connections), and direct steering of these features flips the response between deceptive and honest in a predictable, controlled manner.

What does deception look like inside a language model? Using per-layer transcoders — the same attribution-graph framework Anthropic employs for internal circuit analysis — this paper builds the first mechanistic map of LLM deception, finding a concentrated dictionary of deception-specific features that dominate the causal graph when the model is lying.

Per-layer transcoders (PLTs) applied to Qwen3-4B construct attribution graphs capturing feature activations and inter-feature dependencies across layers. Comparing honest and deceptive completions on the same underlying facts, the paper identifies deception-related features that (1) activate significantly more strongly in the deceptive regime, (2) exert 2–3× greater causal influence on the output distribution than honest-content features, and (3) when directly steered via activation addition, produce predictable, controlled shifts between deceptive and non-deceptive responses. Honest-content features are correspondingly suppressed in the deceptive completion graph. This mechanistic handle on deception is independent of access to training data or reward signals, and the identified features provide interpretable, targetable circuit components for deception detection and mitigation.

Item 4 · Also notable
04
COLM '26 workshop mech-interp

RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models

RARE: Decoupled Routing + Steering in MoE Token Router Steering Expert 1 Expert 2 Expert 3 representation steer Output which expert how to transform
Figure 1: RARE separates expert routing (which expert activates, teal) from representation steering (how the hidden state is transformed, violet) — enabling independent interpretability and targeted intervention on each component in MoE language models.

Accepted to the Actionable Interpretability Workshop at COLM 2026. Standard MoE forward passes couple expert routing with representation transformation, making it difficult to interpret or steer either in isolation. RARE disentangles the two operations — providing separate handles for "which expert" vs. "how the representation is steered" — enabling cleaner interpretability experiments and targeted interventions in MoE architectures.

Notes

2 entries removed on 2026-09-10 as repeats of earlier reports: 2606.25234 (first covered 2026-08-24), 2607.07903 (first covered 2026-07-12).

← all Research Radar issues · gussand · source