📡 Research Radar

September 16, 2026

Pretraining & Midtraining Safety · AI/LLM Security · Applied Mech Interp
Window: Aug 8 – Sep 16, 2026 · Sources: ACL 2026 · EMNLP 2026 · arXiv cs.CL/cs.LG/cs.CR

5 peer-reviewed 5 preprints 0 forum/blog 10 items
pretraining-safety ACL 2026 long paper

Lei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song, Tsung-Yi Ho, Pin-Yu Chen, Yaoqing Yang · ACL 2026 main · arXiv 2506.05346

Research on guardrail collapse has obsessed over fine-tuning methods — SFT vs. RLHF vs. DPO — while ignoring the upstream alignment dataset itself. This ACL long paper shows that dataset similarity is the missing variable: if your safety alignment data looks like the fine-tuning data, your guardrails fail regardless of how carefully you fine-tune.

Dataset Similarity (alignment ↔ fine-tuning task) Harmfulness Score 0 100% −10.33% harm ↓ collapse ↑ Alignment Dataset Similarity → Guardrail Collapse

Figure 3: Dataset similarity vs. post-fine-tuning harmfulness — high similarity predicts guardrail collapse; low-similarity alignment data reduces harmfulness by up to 10.33%.

The authors measure representation similarity (CKA-style) between models' upstream safety alignment datasets and downstream fine-tuning task datasets, then correlate this metric with post-fine-tuning harmfulness scores on jailbreak benchmarks. High alignment-task similarity significantly weakens guardrails; choosing alignment data with low similarity to foreseeable fine-tuning domains reduces harmfulness by up to 10.33% while maintaining utility. The finding reframes guardrail collapse as an alignment-data design problem rather than purely a fine-tuning-method problem: providers should diversify their alignment datasets to reduce overlap with likely downstream task distributions.

Items 4–10
refusal-defense AI-security preprint · Sep 14, 2026

Aashiq Muhamed, Mona T. Diab, Virginia Smith (CMU) · arXiv 2609.16204 · September 14, 2026

Attacker estimates dir. DDO: inject decoy signal into MLP neurons ablates DECOY (harmless) True refusal intact ✓ No base-model retraining required · post-hoc weight edit

DDO corrupts the attacker's contrastive estimator by injecting a decoy; the true refusal circuitry survives. No retraining needed.

Abliteration attacks use contrastive estimators to extract and project out the refusal direction. DDO defeats this by injecting a high-magnitude, nonlinear decoy signal into MLP neurons, corrupting the estimator so it extracts the decoy instead of the true safety mechanism. The true refusal circuitry remains intact; ablation removes only the harmless decoy. The defense is post-hoc and requires no base-model retraining — applicable to already-deployed open-weight models.

safety-decay AI-security ACL 2026 Findings

Rui Zhang, Hongwei Li, Yun Shen, Xinyue Shen, Wenbo Jiang, Guowen Xu, Yang Liu, Michael Backes, Yang Zhang · ACL 2026 Findings · arXiv 2604.07754

SFT-1 SFT-2 SFT-3 SFT-4 DPO ★ ORPO ★ misalign realign HIGHEST ★ BEST ★

Method × safety-direction matrix: ORPO = highest misalignment; DPO = best realignment (utility cost). 4 LLMs × 6 methods. ACL 2026 Findings.

Systematically evaluates four SFT and two preference-optimization methods (ORPO, DPO) across four safety-aligned LLMs for their ability to misalign and re-align. ORPO is most effective for misalignment; DPO most reliably restores alignment but at significant utility cost. Provides an attack/defense method selection matrix directly useful for open-weight model providers.

mech-interp AI-security preprint · Jul 2026

arXiv 2507.00665 · July 1, 2026

Reward Model activations SAE feature ranking chosen vs rejected Target poison Denoising ↓ safety ↑ safety minimal data modification · no general-chat reward degradation

SAE trained on reward model activations; feature-level safety signals drive targeted poisoning or denoising with minimal collateral effect.

Trains a Sparse Autoencoder on reward model activations and ranks features by activation difference between chosen and rejected responses to identify safety-salient latent dimensions. Feature-level signals drive targeted data poisoning (degrading safety alignment) and denoising (enhancing it) with minimal data modification and no general-chat reward degradation — showing reward model safety is concentrated in localizable, interpretable SAE features.

6 entries removed on 2026-09-29 as repeats of earlier reports: 2606.19168 (first covered 2026-08-28), 2609.01455 (first covered 2026-09-03), 2609.00051 (first covered 2026-09-03), 2608.18093 (first covered 2026-08-29), 2606.22686 (first covered 2026-08-18), 2605.26526 (first covered 2026-08-28).

← all Research Radar issues · gussand · source