Nine items, not ten. Every higher-ranked candidate the sweep surfaced (Deep Ignorance, Beyond Safe Data, When Safety Routing Breaks, Fool's Gold, Refusal Aliases, Scalable Circuit Learning, Beyond Shallow Alignment, Representational Alignment) had already been covered by an earlier daily and is excluded by the ledger check. Nothing new appeared in the window on pretraining-time interventions or fine-tuning durability specifically.
01 of 09
Apoorva Upadhyaya, Sandipan Sikdar · arXiv, August 30, 2026
refusal-mechanismmech-interpSAEmultilingualpreprint
Refusal is supposed to be a property of the model, not of the language it happens to be speaking. This paper finds that inside three instruction-tuned models the SAE features that carry refusal share directions with the features that carry "which language is this" — and that you cannot ablate one without moving the other.
Residual-stream SAEs at every layer of Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct and Gemma-2-9B-IT are contrasted on harmful vs. harmless prompts in eight typologically diverse languages to isolate safety-relevant features per language and layer. Their location is architecture-dependent, and within each model they are geometrically entangled with language-identity features: a large fraction of one language's safety features are shared with other languages, with sharing that varies by depth and language pair. The causal test is the point — ablating a language's safety features raises harmful-compliance rates and shifts the output language, so the refusal signal and the language signal ride on overlapping directions, and a feature-level intervention that touches only safety is not on offer.
02 of 09
Thu-Hien Trinh-Thi, Hai-Yen Vong, Thanh-Ha Ung-Dung, Tram Ho (University of Science, VNU-HCM) · arXiv, September 2026
open-weight safety evaljailbreakbenchmarkpreprint
Refusal-rate benchmarks report one number per model. TIER asks a different question: as the threat in a prompt becomes less explicit, does an open-weight model flip from refusing to complying, or does it slide — and where along the slide do models diverge?
Prompts are organized by four risk domains and four threat-implicitness levels — explicit request, contextually embedded, indirect phrasing, full jailbreak framing — and each response is scored on a six-label behavior scale from hard refusal through hedging and partial compliance to full compliance, by two independent LLM judges. Across six open-weight models, behavior degrades gradually across levels; contextual prompts yield the most heterogeneous behavior between models, while the jailbreak level exposes the largest robustness gap. The graded scale makes over-refusal and partial leakage visible where a refusal rate hides them.
03 of 09
University of Science and Technology of China & Zhejiang University · arXiv, August 28, 2026
refusal-mechanismSAE steeringmech-interppreprint
Amplifying a refusal feature works on plain harmful prompts and stops working the moment the request is wrapped in a role-play or formatting task. REINS shows why — the harmful-continuation features are still firing — and fixes it by steering both sides at once.
The authors build GUISE (Generalized Undercover Instruction Safety Evaluation), harmful requests hidden inside complex wrappers, on which existing single-direction SAE steering no longer yields refusals because the harmful-path features stay active and win. REINS suppresses those features and amplifies safe-refusal features in the same feature space. On GUISE it markedly reduces harmful responses and raises genuine refusals while largely preserving general-capability scores; the baselines either intervene too weakly or reach an apparent safety by degenerating output.
Items 4–9
04 of 09
Thomas Rivasseau (McGill University) · arXiv, September 2026
jailbreakfrontier modelspreprint
Cipher jailbreaks previously needed fine-tuning-API access to train a model on encrypted harmful Q&A. This paper shows current frontier models from Anthropic, Google and OpenAI acquire arbitrary ciphers from prompting and in-context examples alone, and that once the exchange runs in the learned cipher, alignment is substantially weakened or bypassed — moving cipher attacks from a fine-tuning-API threat model to an ordinary chat threat model.
05 of 09
Tejasvi C. Addagada · arXiv, September 2026
jailbreakmultilingualpreprint
IICL recasts a harmful request as the last missing cell of a data-labeling task. With a deterministic IICL operator and a StrongREJECT-style judge on gemini-2.5-flash and flash-lite (a provider the original IICL study never tested), over 30 HarmBench general-harm and 30 FinProof financial-abuse behaviors in EN/ES/HI/AR: IICL turns near-total refusal into 80–90% compliance on general harm and 97–100% on financial abuse. Forcing the attack into a lower-resource language blunts it rather than compounding the multilingual gap — measure interactions, don't assume additivity.
06 of 09
XiuYu Zhang, Bonan Ruan, Junfeng Fang, An Zhang, Tat-Seng Chua, Zhenkai Liang (NUS, USTC) · arXiv, August 26, 2026
AI-securitySAEinterp-for-securitypreprint
Adapts the Linux Security Modules separation to LLM serving: a pluggable backend (SAE, transcoder or task-fitted dense probe) emits calibrated per-request evidence, a versioned policy evaluates rules, and a separate gate enforces. Prototyped on Transformers and vLLM. On Qwen3-4B, LMSM-Checkpoint cuts HarmBench attack success from 39.20% to 3.32% while XSTest false refusals rise only from 2.40% to 4.40%, at 98.14% of throughput.
07 of 09
Wissam Antoun et al. · arXiv, September 2026
mech-interpbackdoorSAEpreprint
A controlled backdoor — fixed trigger sequences make 1B and 8B models continue English prompts in French or German — is studied with SAEs across layers and components, contrasting triggered prompts with translation and pretraining controls. SAE features separate triggered prompts from controls with near-perfect F1, but the detecting features do not control the switched behavior: ablating them does not reliably remove the backdoor. Detection and removal need different feature sets.
08 of 09
Yizhe Zeng et al. · arXiv, August 31, 2026
mech-interpbackdoorSAEpreprint
A 2×2 design (clean vs. poisoned model × clean vs. triggered input) traces backdoor-induced logit shifts to high-contribution SAE features and sorts them into interaction, suppressed, mixed and weight-modified roles. Dirty-label and clean-label attacks populate these roles differently, which is why a defense built for one paradigm misses the other — each existing defense implicitly targets one role.
09 of 09
Hoejoon Kwon et al. · arXiv, August 31, 2026
mech-interpsteeringpreprint
Frames safety steering as two coupled decisions — when to intervene and how generation proceeds afterwards — and couples a selective trigger with refusal-anchored constructive redirection in one pass. On Llama-3.1 and Qwen2.5 it preserves benign utility while increasing constructive safe-completion behavior, with the largest gains on models that otherwise emit short refusals.
Notes
- Refusal mechanisms dominate: #1 (refusal features entangled with language), #3 (two-sided SAE steering), #9 (safe-completion steering). Flag #1 for the weekly.
- Two in-window jailbreak results (#4 cipher-in-context, #5 IICL) both show alignment failing under a representational shift the model performs itself.
- #6–#8 use SAEs/probes as a serving-layer control, a forensic tool, and a defense taxonomy respectively.
- This edition replaces an earlier Sep 10 version in which 6 of 10 entries repeated earlier dailies; the repository was retroactively deduplicated the same day.