TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
TamperBench curates weight-space fine-tuning attacks, latent-space representation attacks, and alignment-stage defenses into one evaluation harness. Each of 21 models — including defense-augmented variants — is run against all 9 threats with hyperparameter sweeps per model–attack pair, using standardized safety and capability metrics. The result is the most comprehensive empirical survey of how alignment defenses hold under adversarial pressure.
Key findings: jailbreak-tuning dominates across architectures; base vs. post-trained variants exhibit opposite tamper-resistance trends in Llama-3 and Qwen3 (post-training is not always the safer regime); and no current alignment-stage defense survives the full attack sweep. Code is public.
Attack success rate by threat category across 21 open-weight LLMs; jailbreak-tuning consistently achieves highest ASR (≈82%) — current alignment-stage defenses offer no robust resistance to the full attack sweep.