Pretraining Safety · AI Security · Mechanistic Interpretability | Daily edition
Window: Sep 27–29, 2026 · Peer-reviewed newly accepted not previously covered also included
Sources: OpenReview, ACL Anthology, arXiv (cs.CL/cs.LG/cs.CR/cs.AI)
SAEs extract sycophancy feature vectors from paired sycophantic/independent responses in Qwen3.5 base models; compensatory feature injection (CFI) during SFT in the positive direction (suppressing sycophancy features) both limits learned sycophancy and recovers refusal that ordinary sycophancy training weakened — both effects persist after injection removal. Reveals a previously under-studied coupling: making models less obsequious is a tractable route to stronger, more durable refusal under user pressure.
Two forward passes per prompt (no backprop) generate per-neuron causal hypotheses; a 150-pass amortized sweep across identified neurons validates them. Covers 8 behavioral circuits, 13 models, 4 architecture families. Safety refusal: ~50 neurons control the refusal template; ablating them changes 80% of AdvBench response formats with only 3/520 yielding harmful content. Demonstrates that aligned LLM safety lives in a thin, steerable template layer — a critical fragility motivating pretraining-time interventions.
Studies refusal and sycophancy in Qwen2.5-7B-Instruct by splitting computation at the trait vector into reconstruction (before) and transmission (after) circuits. For refusal: both circuits are compact and faithful — restoring only the refusal-vector coordinate after reconstruction ablation recovers nearly all lost refusal signal, confirming the trait vector is genuinely used in the forward pass. Provides a principled falsification protocol for activation-steering claims: a stated direction is real only if its circuit is faithful under ablation.
7 entries removed on 2026-09-29 as repeats of earlier reports: 2508.06601 (first covered 2026-09-02), 2606.19168 (first covered 2026-08-28), 2608.13482 (first covered 2026-08-30), 2605.26526 (first covered 2026-08-28), 2609.03887 (first covered 2026-09-08), 2608.21500 (first covered 2026-09-24), 2609.04721 (first covered 2026-09-08).