RESEARCH RADAR
Daily · October 8, 2026
1 peer-reviewed · 9 preprints · 0 forum/blog
AI Safety · Alignment · Mech Interp
01
COLM 2026 mech-interp AI security

Component and Dimension Sparsity in Transformer Refusal Mechanisms

Vincent Siu, Glenn Grant-Richards, Vlad Pavlovich, Yizhou Sun, Dawn Song, Chenguang Wang — COLM 2026 (accepted)

The sole peer-reviewed paper this cycle and the most deployable mech-interp result: refusal is not diffusely distributed across a transformer but is carried by a small, identifiable set of components and dimensions—meaning precise interventions are possible, and broad refusal-removal attacks are likely hitting noise.

Component Sparsity → Dimension Sparsity Component level 28% of components 88–101% of steering effect Dimension level (within components) ~50% of dimensions 85–98% of component-level baseline
Figure 1: Two-level sparsity — 28–48% of upstream components carry 88–101% of the refusal steering effect; within those components, ~50% of residual-stream dimensions retain 85–98% of that effect. Consistent with a privileged basis structure.

Siu et al. apply activation steering across four open-weight models and perform ablation sweeps at two granularities: which attention/MLP components are necessary, and which residual-stream dimensions within those components matter. Finding that 28–48% of upstream components account for 88–101% of the steering effect, and roughly half of dimensions within those components retain 85–98% of the component-level baseline, has two direct implications. First, targeted interventions can disable refusal by hitting fewer than half of upstream units. Second, coarse refusal-removal attacks that operate on full-dimension mean-difference vectors are likely coupling to the non-causal dimensions, explaining their brittleness under distribution shift.

02
mech-interp alignment preprint

SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

Miao Yu, Hao Huang, Lu Yuan, Yunpeng Li, Kun Wang, Zuming Jiang — arXiv preprint, 7 Oct 2026

Refusal circuits are latent before any alignment training, but are weak—and they continuously restructure rather than merely strengthen as alignment proceeds. This has direct implications for why fine-tuning attacks erode safety so easily, and for how Safety Circuit Alignment could fix it.

Circuit Overlap Across 20 Alignment Checkpoints 100% 88% 76% 70% 0 10 20 ckpts 98% (ckpt 0) 78% (ckpt 20) 26% — independent runs Y: circuit overlap | X: alignment checkpoint
Figure 2: Adjacent-checkpoint refusal circuit overlap falls from 98% to 78% across 20 alignment steps (teal), while independently extracted circuits share only 26% overlap (red dashed), showing alignment continuously restructures rather than merely amplifying the safety circuit.

SafeEvo formulates circuit extraction as differentiable mask optimization, isolating the sparse subgraph that carries full refusal ability. On Llama-3-8B, Qwen-2.5-7B, and Mistral-7B, ablating the "weak refusal circuit" in the base model raises attack success rates to 32.7–48.4%—confirming refusal is present but unreliable before alignment. Tracking circuits across checkpoints (warm-started extraction) reveals 98→78% overlap decay over 20 steps; independent runs share only 26%, demonstrating circuit non-uniqueness. Safety Circuit Alignment (SCA) confines gradient updates to the discovered refusal circuit, cutting the alignment tax while preserving safety gains.

03
AI security mech-interp preprint

Backdooring Sparse Autoencoders

Enrico Ahlers, Daniel Passon, Tobias Kiecker, Eik Reichmann, Lars Grunske — arXiv preprint (cs.CR), 5 Oct 2026

As SAEs move from interpretability aids to active steering components, a trojanized SAE decoder can silently hijack an unmodified LLM—without touching the model itself and without tripping standard SAEBench quality metrics.

Backdoored SAE Supply-Chain Attack LLM (unchanged) SAE (decoder modified) Malicious output Trigger in prompt → attack fires SAEBench quality metrics: near-unchanged (attack evades standard detection)
Figure 3: Only the SAE decoder is modified; the LLM and SAE encoder are untouched. On trigger, the poisoned reconstruction drives the model to produce attacker-chosen outputs while reconstruction loss and feature-density scores remain nearly normal.

The attack restricts modifications to the SAE decoder, leaving the LM weights and SAE encoder intact. In a code-generation case study across three LLMs and a wide range of insertion layers, the backdoored SAE achieves high rates of unsolicited malicious code insertion and reliable trigger-dependent activation, while reconstruction loss and feature density—the standard SAEBench metrics—show only small changes. The threat model covers any setting where a third party distributes a trained SAE as a drop-in steering component, and the paper concludes that SAEs intended for in-context steering should be cryptographically signed and reproducibly audited rather than treated as benign auxiliary artifacts.

Items 4 – 10 · Also notable
04
AI security preprint

APEX: Active Protection at Execution Boundaries for LLM Agents

Zheng et al. — arXiv (cs.CR), 3 Oct 2026

APEX: Execution-Boundary Defense Untrusted env (tool outputs, injected content) APEX boundary WRAP: authorized? PLANT: endorsed? Safe action 0% ASR on 5/6 benchmarks · 0.56% on 6th
Figure 4: APEX moves defense from pattern matching to the execution boundary, applying authorization (WRAP) and deception-probe (PLANT) checks before any agent action fires.

APEX defends against indirect prompt injection by checking whether a proposed action is authorized and whether the runtime information reaching it was endorsed by the original task—without looking for attack patterns. Against 13 baselines across six benchmarks, it achieves 0% attack success on five of six and 0.56% on the sixth, holding under adaptive attacks across all capability-unit types.


05
AI security alignment preprint

Does On-Policy Distillation for Safety Pose Backdoor Risks?

Jian Luo et al. — arXiv (cs.CR), 6 Oct 2026

Backdoor Transfer via On-Policy Distillation Teacher LLM (safety-aligned + On-Policy Distillation 3% poison rate Student ASR 70% Lazy Defense (KL clipping) delays transfer at low poison rates
Figure 5: A backdoored safety-aligned teacher transfers to a clean student via on-policy distillation at just 3% poisoning rate (70% ASR), amplified by more training epochs.

A safety-aligned teacher with 3% poisoned training data transfers backdoor behavior to a clean student via on-policy distillation at 70% attack success rate; even 10 poisoned samples reach 67% ASR after 16 epochs. Top-k KL estimation accelerates transfer relative to sampled-token KL. The proposed Lazy Defense clips KL rewards to slow student updates and delay backdoor transfer at low poisoning rates.


06
mech-interp AI security preprint

Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild

Elad David, Max Fomin — arXiv, 1 Oct 2026

Classification Suffix → Better OOD Probe Generalization User turn → probe read No suffix (baseline) OOD AUC: lower User turn + classify suffix → probe Content-free label ≈ real label OOD AUC: +up to ~4 pts 13 safety benchmarks · 3 model families · LODO eval
Figure 6: Appending a classification instruction after the user turn consistently improves leave-one-dataset-out AUC of activation-based malicious-input probes by up to ~4 points; format matters more than label content.

Appending a short classification suffix after the user turn at probe-read time improves out-of-distribution generalization of activation-based malicious-input detectors by up to ~4 AUC points under leave-one-dataset-out evaluation across 13 safety benchmarks and three model families. Format (classification instruction) matters more than content (a content-free label suffix matches a real label suffix); deployed via KV-cache fork, it is a near-zero-cost drop-in for any activation-probe monitor.


07
AI security preprint

Secure Speculative Decoding for Large Language Models

Yichi Zhang, Zhiqi Wang, Neil Gong, Yuchen Yang — arXiv (cs.CR), 6 Oct 2026

SecureSD: Security–Efficiency Trade-off Decoding speedup → Attack success rate → Baseline SD (high ASR) SecureSD (99.8% speedup kept) −92.4% ASR
Figure 7: SecureSD tightens early-token verification in the draft model, collapsing attack success rate by up to 92.4% while keeping 99.8% of the decoding speedup.

SecureSD identifies a security-utility asymmetry in lossy speculative decoding—efficiency gains cause disproportionate security loss—and fixes it by applying stricter verification at early draft-token positions where the vulnerability originates. Across five benchmarks, it reduces jailbreak and prompt injection attack success rates by up to 92.4% while retaining 99.8% of the decoding speedup and 98.4% of utility.


08
AI security preprint

Localize-and-Detect: Auditing Task-Level Poisoning in Instruction-Tuned Models

Luze Sun, Cristina Nita-Rotaru, Alina Oprea — arXiv (cs.CR), 2 Oct 2026

Localize-and-Detect: Two-Stage Black-Box Auditor Stage 1: Localize Rank tasks by distribution divergence (base vs FT model) Stage 2: Detect Search top tasks for biased content in FT responses Black-box only (no internals) · 216 poisoned models · 2 LLM families
Figure 8: Two-stage black-box auditor for task-conditioned poisoning: distribution-divergence ranking (Stage 1) feeds into biased-content detection (Stage 2), with no access to model weights required.

A two-stage black-box auditor for task-level (non-trigger) poisoning: Stage 1 ranks tasks by output-distribution divergence between fine-tuned and base models; Stage 2 finds biased content in the fine-tuned model's top-ranked task responses. Evaluated on 216 poisoned models across two LLM families, white-box signals do not outperform the black-box approach, making the auditor deployable without model access.


09
alignment preprint

PyINE: A Framework for Scalable Elicitation and Oversight via Code Execution

Pierre-Luc St-Charles, Yoshua Bengio et al. — arXiv (cs.AI), 3 Oct 2026

PyINE: Code Execution as Oversight Substrate Python programs (~1M exec traces) LLM code variants (500K+ matched) Oversight eval (probes, judges) Deceptive-cue failure: model errs when human-readable cues conflict with exec result
Figure 9: PyINE uses Python execution traces as verifiable ground truth for oversight research, exposing a deceptive-cue failure mode where RL-trained models err when human-readable labels conflict with actual program behavior.

PyINE provides ~1M deterministic Python execution traces and 500K+ LLM-generated code variants as a scalable, mechanically generated oversight substrate (Yoshua Bengio et al.). RL-trained models improve substantially at execution-outcome prediction but still fail when human-readable cues conflict with actual program behavior—a concrete, measurable instance of the deceptive-cue problem in scalable oversight.


10
AI security preprint

From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents

Juanyang Xu, Zheng Wang, Xingyu Zhao, Siddartha Khastgir, Andi Zhang — arXiv, 7 Oct 2026

Unifying Jailbreak Objectives via Gradient Identity Maximize E[harmfulness] (expected harm over outputs) ≡ Maximize P(harmful target) (likelihood of target output) → OPUR: harmfulness-reweighted sampling guides input optimization Validated on InjecAgent-style indirect-injection benchmarks
Figure 10: The gradient of expected harmfulness equals the expected gradient of log-likelihood under harmfulness-reweighted outputs—unifying two jailbreak paradigms and motivating OPUR, which samples from the reweighted distribution to guide input optimization.

Proves by gradient identity that maximizing expected harmfulness and maximizing the likelihood of a harmful target output are equivalent jailbreak objectives, unifying two previously separate paradigms. OPUR exploits this identity by sampling from a harmfulness-reweighted output distribution to guide likelihood-based input optimization, demonstrating improved effectiveness on InjecAgent-style indirect-injection benchmarks against LLM agents.

← all Research Radar issues · gussand · source