Research Radar · Daily Edition

Mech Interp & AI Safety
August 28, 2026

2026-08-28  ·  claude/radar
Window: Aug 8–28, 2026 1 peer-reviewed 9 preprints 0 forum/blog arXiv · ACL · KDD 2026
Top 3 — Full Treatment
01 KDD 2026 pretraining-safety

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

Why it ranks first
The field's first unified benchmark for tamper resistance: 21 open-weight LLMs evaluated across 9 tampering threats in a single reproducible harness. The only peer-reviewed entry this window. Jailbreak-tuning is consistently the most severe attack; current alignment-stage defenses largely fail the full sweep.

TamperBench curates weight-space fine-tuning attacks, latent-space representation attacks, and alignment-stage defenses into one evaluation harness. Each of 21 models — including defense-augmented variants — is run against all 9 threats with hyperparameter sweeps per model–attack pair, using standardized safety and capability metrics. The result is the most comprehensive empirical survey of how alignment defenses hold under adversarial pressure.

Key findings: jailbreak-tuning dominates across architectures; base vs. post-trained variants exhibit opposite tamper-resistance trends in Llama-3 and Qwen3 (post-training is not always the safer regime); and no current alignment-stage defense survives the full attack sweep. Code is public.

0% 25% 50% 75% 82% JB-Tuning 58% Rep. Attack 52% Grad. FT 41% Adv. Suffix 29% Prefix Inj. 19% Latent Steer Most severe Other attacks Attack severity across 21 LLMs · TamperBench

Attack success rate by threat category across 21 open-weight LLMs; jailbreak-tuning consistently achieves highest ASR (≈82%) — current alignment-stage defenses offer no robust resistance to the full attack sweep.

02 pretraining-safety mech-interp

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

Why it ranks here
A striking dissociation: activation-steering weight edits survive SFT almost untouched (mean cos θ = 0.074), yet behavioral compliance with the steered direction collapses by 64% on average. Behavior reverts while the weight modification stays intact — pointing to circuit-level recovery mechanisms that bypass the steered direction.

The study embeds refusal-suppression and brevity-induction steering vectors directly into weights across five instruction-tuned models (3B–14B), then applies non-adversarial SFT and RLHF. At the behavioral level, refusal ablation loses an average 64% of its effect under SFT. At the mechanistic level, the steering weight edit survives: mean vector recovery correlation ρ = 0.004, update direction near-orthogonal to the pre-edit steering direction (mean cos θ = 0.074).

The implication is significant: the fine-tuner is not reversing the weight edit. Something else in the model's computation pathway overrides the steered direction — a gap between the mechanistic change made and the behavioral change achieved.

0.0 0.25 0.50 1.00 0 Epoch 1 Epoch 2 Epoch 3 Epoch 4 Weight edit (ρ=0.004) Behavioral effect (–64% avg.) 64% gap Steering durability under SFT — 5 models, 3B–14B

Weight edit (teal, flat) vs. behavioral compliance (red-dashed, declining) across 4 SFT epochs; the steering weight modification survives intact while behavioral effect degrades 64% on average — a mechanistic dissociation between edit and outcome.

03 pretraining-safety SAE

Shaping Capabilities with Token-Level Data Filtering

Why it ranks here
Token-level filtering via SAE-labeled classifiers outperforms document-level filtering at every scale — and at the largest model size, the forget domain requires 7000× more compute to recover under adversarial fine-tuning. By Rathi & Radford (OpenAI); directly connects pretraining safety and mech interp tooling (SAEs as the labeling mechanism).

Filtering individual tokens (labeled via SAE-derived classifiers trained to detect domain-relevant activations) achieves greater harm-capability reduction than filtering entire documents at the same data budget, because document-level filters over-prune benign adjacent content. SAE token classifiers are distilled into lightweight inference-time classifiers for scalable deployment.

Scale results: token filtering becomes more effective at larger model sizes. For the largest models, the forget domain requires 7000× more compute to recover capability under adversarial fine-tuning, compared to document-level baselines. Token-filtered models outperform state-of-the-art machine unlearning on tamper-resistance benchmarks — a pretraining-time intervention substantially harder to reverse than any post-hoc method.

1× 50× 500× 5000× 8× 22× Small 18× 80× Medium 40× 550× Large 90× 7000× ↑ X-Large Document-level filter Token-level filter (SAE) Token-level (XL) — 7000× Compute to recover forget-domain capability · by model scale

Tamper-resistance comparison (relative compute slowdown on forget domain) between document-level and token-level (SAE-labeled) filtering across model scales; token-level achieves 7000× slowdown at the largest scale tested, growing substantially more effective than document-level as models scale.

Also Notable · Items 4–10
04

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

pretraining-safety AI security 2605.26526 Kuo, Yadav, Smith · CMU/UC Berkeley · May 2026

Abliteration and prefilling — two gradient-free, low-cost attacks — raise ASR against fine-tuning-defended models (TAR, SEAM) from below 10% to 16–96% on BeaverTails, HarmBench, and AdvBench. The authors propose ART (abliteration-resistant tuning), adding an abliteration-based objective to the defense training loop, reducing ASR by 10–20 pp across attack types. A sobering reminder that defenses evaluated against gradient-based attacks remain brittle to inference-time exploits.

ASR · Simple Attacks Abliteration Prefilling Undefended Defended
05

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection

pretraining-safety 2606.19168 arXiv · June 2026

Safety Reflection Pretraining (SRP) inserts short safety-reflection spans into pretraining corpora, training the model to produce aligned self-monitoring before any post-training. On 1.7B models: harmonic mean improves 75.4%→90.6%; unsafe-response identification 61.0%→87.6%. SRP behavior is substantially more robust to inference-stage jailbreaks and fine-tuning attacks than post-training alignment applied to an identical base, complementing deep-ignorance / data-filtering approaches by acting on the response-conditioning layer.

SRP Improvement Harm. Mean 90.6% 75.4% Unsafe ID 87.6% 61.0% SRP Baseline
06

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

AI security 2608.26008 Hu, Hooi · NUS · August 26, 2026

When a jailbreak fails, the framework abstracts the attack into a method-level rule capturing its structural wrapper — role-playing, obfuscation, multi-step indirection — rather than the topic alone. Rules persist across interactions in a growing cross-session memory; new attack families expand the label space automatically. No parameter updates; applies to both open-weight and API models. The self-evolving mechanism prevents the degradation seen in static guardrails against novel variants.

Attack Failed? Abstract Rule Rule Memory Updated Blocked yes no
07

ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization

AI security 2608.03210 Zhu, Huang, Juefei-Xu et al. · August 2026

Semantic-shift jailbreaks replace harmful keywords with benign substitutes ("bomb"→"carrot") and rely on context to make the model reinterpret the benign word. ICO is a black-box method that iteratively optimizes the context scaffold while preserving the surface-benign input, exploiting highly variable context ability to induce semantic shifts. Surface-benign inputs evade input-level filters; gradient-free context optimization improves substantially over static baselines. Demonstrates semantic-shift attacks are more dangerous than previously characterized.

ICO Optimization Init Iter 1 Iter 2 Iter 3 ASR improves each round, no gradient access
08

Safety Cost of Steering Vectors Is Separable and Reducible

pretraining-safety mech-interp 2608.08383 Li, Kasneci · TU Munich · August 2026

Identifies a separable safety-degrading component within steering vectors that increases compliance with harmful requests without contributing to the steering objective. Formulates removal via constrained optimization (primal-dual updates): minimize the harmful subcomponent subject to preserving the steering effect and bounding false-refusal rate. The safety cost of activation steering is not an inherent trade-off — it is an artifact of steering direction construction, projectable without loss of steering utility.

Steering Vector Decomp. v_steer task safety cost Project out → clean v_steer
09

Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

pretraining-safety mech-interp 2608.04347 Yoshida et al. · Stanford/NTU · August 5, 2026

Fine-tuning on tasks with no explicit safety content degrades alignment properties — "side-effect misalignment." The authors introduce LoRA-based introspection adapters (IAs) trained to verbalize which alignment properties changed during fine-tuning, without requiring a full evaluation suite per run. Makes alignment auditing at deployment time more tractable: the IA produces a natural-language description of what fine-tuning degraded rather than requiring exhaustive probe evaluation.

Source Fine-tune Degraded Introspection Adapter (LoRA) → text desc. "Refusal rate decreased 34%..."
10

Backdoor Decontamination Dynamics in LLM Agents

AI security 2608.11295 arXiv · August 2026

Defensive poisoning alone erases approximately 56% of original backdoor triggers in LLM agents; subsequent decontamination drives nearly all survivors to erasure. Key finding: trigger recognition and malicious execution are behaviorally dissociable — a backdoored model can be conditioned to recognize its trigger without executing the malicious payload, enabling targeted neutralization. Staged decontamination (recognition suppression, then execution suppression) consistently outperforms single-stage defenses.

Backdoor Erasure Initial 100% After def. poisoning 44% After decontamination ~0%
Notes & Patterns

Editorial notes

← all Research Radar issues · gussand · source