📡 Research Radar

Daily Radar — September 29, 2026

Pretraining Safety · AI Security · Mechanistic Interpretability | Daily edition

Window: Sep 27–29, 2026 · Peer-reviewed newly accepted not previously covered also included
Sources: OpenReview, ACL Anthology, arXiv (cs.CL/cs.LG/cs.CR/cs.AI)

3 peer-reviewed 7 preprints 0 forum/blog 10 items



Items 4–10 · Mid-Tier
07 / 10 — Mech Interp · Alignment 🆕 Sep 28

Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability

mech-interp · SAE alignment preprint
Sycophancy RLHF training SAE sycophancy features (Qwen3.5) CFI (+) suppresses syco → recovers refusal Sycophancy training is a hidden threat to refusal robustness; SAE-based CFI mitigates it
Figure 7: Causal chain — SAE extracts sycophancy features; CFI positive injection during SFT suppresses sycophancy and recovers refusal that RLHF sycophancy training degraded.

SAEs extract sycophancy feature vectors from paired sycophantic/independent responses in Qwen3.5 base models; compensatory feature injection (CFI) during SFT in the positive direction (suppressing sycophancy features) both limits learned sycophancy and recovers refusal that ordinary sycophancy training weakened — both effects persist after injection removal. Reveals a previously under-studied coupling: making models less obsequious is a tractable route to stronger, more durable refusal under user pressure.

08 / 10 — Applied Mech Interp · AI Security

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs

mech-interp AI security preprint
Pass 1: normal Pass 2: perturbed Rank FFN neurons by Δ activation ~50 neurons (0.014%) → 80% format changes
Figure 8: Two-pass perturbation probing isolates ~50 FFN neurons (0.014% of the network) that control 80% of safety-refusal format changes on AdvBench, with near-zero harmful compliance.

Two forward passes per prompt (no backprop) generate per-neuron causal hypotheses; a 150-pass amortized sweep across identified neurons validates them. Covers 8 behavioral circuits, 13 models, 4 architecture families. Safety refusal: ~50 neurons control the refusal template; ablating them changes 80% of AdvBench response formats with only 3/520 yielding harmful content. Demonstrates that aligned LLM safety lives in a thin, steerable template layer — a critical fragility motivating pretraining-time interventions.

10 / 10 — Applied Mech Interp · Alignment

From Trait Vectors to Circuits: Tracing Refusal and Sycophancy Through Language Models

mech-interp alignment preprint
before vector Reconstruction circuit (compact + faithful) Refusal trait vector Transmission circuit (after vector) Restoring the refusal coordinate after reconstruction ablation recovers ~all of the refusal signal
Figure 10: Refusal circuit split in Qwen2.5-7B — compact reconstruction and transmission subgraphs; restoring the vector coordinate recovers nearly all ablated refusal signal.

Studies refusal and sycophancy in Qwen2.5-7B-Instruct by splitting computation at the trait vector into reconstruction (before) and transmission (after) circuits. For refusal: both circuits are compact and faithful — restoring only the refusal-vector coordinate after reconstruction ablation recovers nearly all lost refusal signal, confirming the trait vector is genuinely used in the forward pass. Provides a principled falsification protocol for activation-steering claims: a stated direction is real only if its circuit is faithful under ablation.

7 entries removed on 2026-09-29 as repeats of earlier reports: 2508.06601 (first covered 2026-09-02), 2606.19168 (first covered 2026-08-28), 2608.13482 (first covered 2026-08-30), 2605.26526 (first covered 2026-08-28), 2609.03887 (first covered 2026-09-08), 2608.21500 (first covered 2026-09-24), 2609.04721 (first covered 2026-09-08).

← all Research Radar issues · gussand · source