Top 3 · Full treatment
01 · Pretraining Safety
Yitong Guo, Xiaoyi Chen, Siyuan Zhang, Xiaofeng Wang, Haixu Tang · arXiv:2609.01455 · September 1, 2026
First mechanistic (Fisher-geometric) explanation of why even capability-only fine-tuning reliably breaks safety alignment — the why behind a well-known empirical failure mode, with direct implications for open-weight deployment.
Alignment flattens the Fisher information geometry for safety-critical weights (low-rank safety Fisher), creating a dormant output-routing pathway through output-side MLP modules. As few as 100 benign fine-tuning examples selectively re-sharpen this pathway — attack success rates jump from ~3% to ~71% while MMLU degrades mildly. The asymmetry is output-localized: safety representations are preserved (explaining why a handful of safety examples can quickly restore refusal), but the routing geometry is disrupted. LoRA and ASAM suppress early collapse by dampening output-side sharpness, but protection degrades at larger fine-tuning budgets.
pretraining-safety
fine-tuning-attack
Fisher-geometry
alignment-fragility
arXiv:2609.01455
03 · AI Security · Multimodal
Tianqi Xiao, Shiyao Cui, Minghao Zhang, Junxiao Yang, Renmiao Chen · arXiv:2609.02082 · September 2, 2026
Identifies and measures a specific multimodal safety gap (cross-modal safety drift) and proposes a training-free fix — directly relevant to deployed multimodal systems.
Cross-modal safety drift: a benign textual query paired with a harmful image bypasses refusal far more often than an explicitly harmful text query. Visually-risky cues receive limited attention in the model's representation and weakly trigger safety-relevant computations. Safety-Awareness Representation Transfer (SRT) is a lightweight, training-free direction-refinement method applied to a frozen MLLM backbone. It projects the model's internal state toward the unsafe-text safety direction when a visually-risky query is detected, substantially reducing compliance with cross-modal harmful requests without degrading benign multimodal performance.
multimodal-safety
cross-modal-drift
training-free
MLLM
arXiv:2609.02082
Items 4–10 · Also notable
Hongli Shen, Shaopeng Fu, Qinbo Zhang, Jian Li, Di Wang (KAUST, Stony Brook) · arXiv:2608.09542 · Aug 10, 2026
AdvSafe trains large reasoning models to internalize attack mechanisms, not pattern-match on jailbreak surface forms. A synthesis agent dynamically crafts diverse jailbreaks; the target LRM learns from both the attack decomposition and the rejection. With 1,000 synthesized samples, AdvSafe-aligned LRMs significantly outperform robustness baselines with near-zero utility degradation.
AI-security
jailbreak
LRM
preprint
Kuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng (UC San Diego) · arXiv:2609.00051 · EMNLP 2026 Findings (peer-reviewed)
Characterizes a three-stage safety circuit: Harmful Detection Heads → Safety Neurons → Refusal Heads. Circuit-guided weight scaling (persistent weight-space intervention, not activation steering) of the identified components improves safety rates under adversarial attacks by 26.5% across six LLMs with only 1.7% accuracy drop on four standard benchmarks.
mech-interp-applied
AI-security
peer-reviewed
EMNLP-2026-Findings
Alireza Lotfi, Subangkar K. Shanto, Imtiaz Karim, Elisa Bertino · arXiv:2608.01558 · ACM AI Leadership Summit 2026, August 2026
Agentic AI safety is a trajectory property, not a per-action one: a sequence of individually-approved actions can violate system invariants in aggregate. Proposes Trajectory Assurance as a security primitive — runtime monitors enforcing invariants over complete action sequences across organizational boundaries. Covers the full agentic attack surface at both single-agent and multi-agent delegation layers.
AI-security
agentic
trajectory-assurance
preprint
Notes
- Deep Ignorance (2508.06601) and Synthetic Persona Pretraining (2608.13482) were covered in the 2026-09-02 report; both remain high-priority references for pretraining-time safety.
- Freshest papers this window: 2609.01455 (Sept 1) and 2609.02082 (Sept 2) — both September 2026 batch.
- The September 2026 arXiv batch (2609.xxxxx) is just opening; expect higher volume through the week.
- Flagged for weekly: rank 1 (Fisher-geometric alignment fragility), rank 5 (circuit-guided weight scaling, EMNLP peer-reviewed), rank 6 (separable steering safety, COLM peer-reviewed), rank 7 (EM attribution via SAE).