Vincent Siu, Glenn Grant-Richards, Vlad Pavlovich, Yizhou Sun, Dawn Song, Chenguang Wang — COLM 2026 (accepted)
The sole peer-reviewed paper this cycle and the most deployable mech-interp result: refusal is not diffusely distributed across a transformer but is carried by a small, identifiable set of components and dimensions—meaning precise interventions are possible, and broad refusal-removal attacks are likely hitting noise.
Figure 1: Two-level sparsity — 28–48% of upstream components carry 88–101% of the refusal steering effect; within those components, ~50% of residual-stream dimensions retain 85–98% of that effect. Consistent with a privileged basis structure.
Siu et al. apply activation steering across four open-weight models and perform ablation sweeps at two granularities: which attention/MLP components are necessary, and which residual-stream dimensions within those components matter. Finding that 28–48% of upstream components account for 88–101% of the steering effect, and roughly half of dimensions within those components retain 85–98% of the component-level baseline, has two direct implications. First, targeted interventions can disable refusal by hitting fewer than half of upstream units. Second, coarse refusal-removal attacks that operate on full-dimension mean-difference vectors are likely coupling to the non-causal dimensions, explaining their brittleness under distribution shift.
Miao Yu, Hao Huang, Lu Yuan, Yunpeng Li, Kun Wang, Zuming Jiang — arXiv preprint, 7 Oct 2026
Refusal circuits are latent before any alignment training, but are weak—and they continuously restructure rather than merely strengthen as alignment proceeds. This has direct implications for why fine-tuning attacks erode safety so easily, and for how Safety Circuit Alignment could fix it.
Figure 2: Adjacent-checkpoint refusal circuit overlap falls from 98% to 78% across 20 alignment steps (teal), while independently extracted circuits share only 26% overlap (red dashed), showing alignment continuously restructures rather than merely amplifying the safety circuit.
SafeEvo formulates circuit extraction as differentiable mask optimization, isolating the sparse subgraph that carries full refusal ability. On Llama-3-8B, Qwen-2.5-7B, and Mistral-7B, ablating the "weak refusal circuit" in the base model raises attack success rates to 32.7–48.4%—confirming refusal is present but unreliable before alignment. Tracking circuits across checkpoints (warm-started extraction) reveals 98→78% overlap decay over 20 steps; independent runs share only 26%, demonstrating circuit non-uniqueness. Safety Circuit Alignment (SCA) confines gradient updates to the discovered refusal circuit, cutting the alignment tax while preserving safety gains.
Enrico Ahlers, Daniel Passon, Tobias Kiecker, Eik Reichmann, Lars Grunske — arXiv preprint (cs.CR), 5 Oct 2026
As SAEs move from interpretability aids to active steering components, a trojanized SAE decoder can silently hijack an unmodified LLM—without touching the model itself and without tripping standard SAEBench quality metrics.
Figure 3: Only the SAE decoder is modified; the LLM and SAE encoder are untouched. On trigger, the poisoned reconstruction drives the model to produce attacker-chosen outputs while reconstruction loss and feature-density scores remain nearly normal.
The attack restricts modifications to the SAE decoder, leaving the LM weights and SAE encoder intact. In a code-generation case study across three LLMs and a wide range of insertion layers, the backdoored SAE achieves high rates of unsolicited malicious code insertion and reliable trigger-dependent activation, while reconstruction loss and feature density—the standard SAEBench metrics—show only small changes. The threat model covers any setting where a third party distributes a trained SAE as a drop-in steering component, and the paper concludes that SAEs intended for in-context steering should be cryptographically signed and reproducibly audited rather than treated as benign auxiliary artifacts.
Figure 4: APEX moves defense from pattern matching to the execution boundary, applying authorization (WRAP) and deception-probe (PLANT) checks before any agent action fires.
APEX defends against indirect prompt injection by checking whether a proposed action is authorized and whether the runtime information reaching it was endorsed by the original task—without looking for attack patterns. Against 13 baselines across six benchmarks, it achieves 0% attack success on five of six and 0.56% on the sixth, holding under adaptive attacks across all capability-unit types.
Figure 5: A backdoored safety-aligned teacher transfers to a clean student via on-policy distillation at just 3% poisoning rate (70% ASR), amplified by more training epochs.
A safety-aligned teacher with 3% poisoned training data transfers backdoor behavior to a clean student via on-policy distillation at 70% attack success rate; even 10 poisoned samples reach 67% ASR after 16 epochs. Top-k KL estimation accelerates transfer relative to sampled-token KL. The proposed Lazy Defense clips KL rewards to slow student updates and delay backdoor transfer at low poisoning rates.
Figure 6: Appending a classification instruction after the user turn consistently improves leave-one-dataset-out AUC of activation-based malicious-input probes by up to ~4 points; format matters more than label content.
Appending a short classification suffix after the user turn at probe-read time improves out-of-distribution generalization of activation-based malicious-input detectors by up to ~4 AUC points under leave-one-dataset-out evaluation across 13 safety benchmarks and three model families. Format (classification instruction) matters more than content (a content-free label suffix matches a real label suffix); deployed via KV-cache fork, it is a near-zero-cost drop-in for any activation-probe monitor.
Yichi Zhang, Zhiqi Wang, Neil Gong, Yuchen Yang — arXiv (cs.CR), 6 Oct 2026
Figure 7: SecureSD tightens early-token verification in the draft model, collapsing attack success rate by up to 92.4% while keeping 99.8% of the decoding speedup.
SecureSD identifies a security-utility asymmetry in lossy speculative decoding—efficiency gains cause disproportionate security loss—and fixes it by applying stricter verification at early draft-token positions where the vulnerability originates. Across five benchmarks, it reduces jailbreak and prompt injection attack success rates by up to 92.4% while retaining 99.8% of the decoding speedup and 98.4% of utility.
Luze Sun, Cristina Nita-Rotaru, Alina Oprea — arXiv (cs.CR), 2 Oct 2026
Figure 8: Two-stage black-box auditor for task-conditioned poisoning: distribution-divergence ranking (Stage 1) feeds into biased-content detection (Stage 2), with no access to model weights required.
A two-stage black-box auditor for task-level (non-trigger) poisoning: Stage 1 ranks tasks by output-distribution divergence between fine-tuned and base models; Stage 2 finds biased content in the fine-tuned model's top-ranked task responses. Evaluated on 216 poisoned models across two LLM families, white-box signals do not outperform the black-box approach, making the auditor deployable without model access.
Pierre-Luc St-Charles, Yoshua Bengio et al. — arXiv (cs.AI), 3 Oct 2026
Figure 9: PyINE uses Python execution traces as verifiable ground truth for oversight research, exposing a deceptive-cue failure mode where RL-trained models err when human-readable labels conflict with actual program behavior.
PyINE provides ~1M deterministic Python execution traces and 500K+ LLM-generated code variants as a scalable, mechanically generated oversight substrate (Yoshua Bengio et al.). RL-trained models improve substantially at execution-outcome prediction but still fail when human-readable cues conflict with actual program behavior—a concrete, measurable instance of the deceptive-cue problem in scalable oversight.
Figure 10: The gradient of expected harmfulness equals the expected gradient of log-likelihood under harmfulness-reweighted outputs—unifying two jailbreak paradigms and motivating OPUR, which samples from the reweighted distribution to guide input optimization.
Proves by gradient identity that maximizing expected harmfulness and maximizing the likelihood of a harmful target output are equivalent jailbreak objectives, unifying two previously separate paradigms. OPUR exploits this identity by sampling from a harmfulness-reweighted output distribution to guide likelihood-based input optimization, demonstrating improved effectiveness on InjecAgent-style indirect-injection benchmarks against LLM agents.