📡 Research Radar — September 10, 2026

Pretraining Safety · AI Security · Mech Interp · Daily edition
Window: Sep 8–10, 2026, plus Aug 26 – Sep 7 papers no earlier report covered · Sources: OpenReview, arXiv (cs.CL/cs.LG/cs.CR/cs.AI), ACL Anthology
Counts: 0 peer-reviewed · 9 preprints · 0 forum/blog

Nine items, not ten. Every higher-ranked candidate the sweep surfaced (Deep Ignorance, Beyond Safe Data, When Safety Routing Breaks, Fool's Gold, Refusal Aliases, Scalable Circuit Learning, Beyond Shallow Alignment, Representational Alignment) had already been covered by an earlier daily and is excluded by the ledger check. Nothing new appeared in the window on pretraining-time interventions or fine-tuning durability specifically.

01 of 09

When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs

refusal-mechanismmech-interpSAEmultilingualpreprint

Refusal is supposed to be a property of the model, not of the language it happens to be speaking. This paper finds that inside three instruction-tuned models the SAE features that carry refusal share directions with the features that carry "which language is this" — and that you cannot ablate one without moving the other.

Share of a language's safety features that are also language-identity features, by layer 00.51.0 earlylayerlate Llama-3.1-8B-InstructQwen2.5-7B-InstructGemma-2-9B-IT Schematic of the paper's finding; exact curves are in the paper (8 languages: EN ZH DE AR VI ID RU HI)

Figure 1 (schematic): safety-relevant SAE features overlap language-identity features throughout the network, with architecture-specific layer profiles; ablating them raises harmful compliance and changes the output language.

Residual-stream SAEs at every layer of Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct and Gemma-2-9B-IT are contrasted on harmful vs. harmless prompts in eight typologically diverse languages to isolate safety-relevant features per language and layer. Their location is architecture-dependent, and within each model they are geometrically entangled with language-identity features: a large fraction of one language's safety features are shared with other languages, with sharing that varies by depth and language pair. The causal test is the point — ablating a language's safety features raises harmful-compliance rates and shifts the output language, so the refusal signal and the language signal ride on overlapping directions, and a feature-level intervention that touches only safety is not on offer.

02 of 09

TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors

open-weight safety evaljailbreakbenchmarkpreprint

Refusal-rate benchmarks report one number per model. TIER asks a different question: as the threat in a prompt becomes less explicit, does an open-weight model flip from refusing to complying, or does it slide — and where along the slide do models diverge?

Behavior mix across threat-implicitness levels (six-label scale, six open-weight LLMs) explicitcontextualindirectjailbreak refuse hedge/partial comply

Figure 1 (schematic): behavior shifts gradually across the four levels rather than flipping at one threshold; the jailbreak level shows the widest gap between models.

Prompts are organized by four risk domains and four threat-implicitness levels — explicit request, contextually embedded, indirect phrasing, full jailbreak framing — and each response is scored on a six-label behavior scale from hard refusal through hedging and partial compliance to full compliance, by two independent LLM judges. Across six open-weight models, behavior degrades gradually across levels; contextual prompts yield the most heterogeneous behavior between models, while the jailbreak level exposes the largest robustness gap. The graded scale makes over-refusal and partial leakage visible where a refusal rate hides them.

03 of 09

REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features

refusal-mechanismSAE steeringmech-interppreprint

Amplifying a refusal feature works on plain harmful prompts and stops working the moment the request is wrapped in a role-play or formatting task. REINS shows why — the harmful-continuation features are still firing — and fixes it by steering both sides at once.

Same SAE feature space, two interventions per token harmful-continuation features ↓ suppress left active by refusal-only steering safe-refusal features ↑ enhance without pushing into collapse + Evaluated on GUISE: harmful prompts inside complex wrappers

Figure 1 (schematic): REINS inhibits harmful-continuation SAE features while enhancing refusal features; refusal-only steering leaves the former active on wrapped prompts.

The authors build GUISE (Generalized Undercover Instruction Safety Evaluation), harmful requests hidden inside complex wrappers, on which existing single-direction SAE steering no longer yields refusals because the harmful-path features stay active and win. REINS suppresses those features and amplifies safe-refusal features in the same feature space. On GUISE it markedly reduces harmful responses and raises genuine refusals while largely preserving general-capability scores; the baselines either intervene too weakly or reach an apparent safety by degenerating output.

Items 4–9

04 of 09

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning

jailbreakfrontier modelspreprint
Teach cipher in-contextprompting only, no fine-tuning Ask in cipherharmful request, encoded Answer in cipheralignment weakened or bypassed →→ Demonstrated against Anthropic, Google and OpenAI frontier models

Figure 1 (schematic): the cipher is acquired from in-context examples; safety behavior that holds in plaintext is weakened once the exchange moves into the cipher.

Cipher jailbreaks previously needed fine-tuning-API access to train a model on encrypted harmful Q&A. This paper shows current frontier models from Anthropic, Google and OpenAI acquire arbitrary ciphers from prompting and in-context examples alone, and that once the exchange runs in the learned cipher, alignment is substantially weakened or bypassed — moving cipher attacks from a fine-tuning-API threat model to an ordinary chat threat model.

05 of 09

Structural Jailbreaks Generalize but Do Not Compound: A Cross-Provider and Multilingual Study of Involuntary In-Context Learning

jailbreakmultilingualpreprint
Compliance rate, gemini-2.5-flash: single-shot baseline vs. IICL 80–100% ENESHIAR red = IICL · grey = baseline (≈0)

Figure 1 (schematic): IICL lifts compliance from near zero to 80–100% in English; the same operator in lower-resource languages is weaker, not stronger.

IICL recasts a harmful request as the last missing cell of a data-labeling task. With a deterministic IICL operator and a StrongREJECT-style judge on gemini-2.5-flash and flash-lite (a provider the original IICL study never tested), over 30 HarmBench general-harm and 30 FinProof financial-abuse behaviors in EN/ES/HI/AR: IICL turns near-total refusal into 80–90% compliance on general harm and 97–100% on financial abuse. Forcing the attack into a lower-resource language blunts it rather than compounding the multilingual gap — measure interactions, don't assume additivity.

06 of 09

LMSM: LLM Security Framework Inspired by Linux Security Modules

AI-securitySAEinterp-for-securitypreprint
Security backendSAE · transcoder · dense probe Versioned policyrules over trusted context Enforcement gateHarmBench ASR 39.2% → 3.32% →→ Qwen3-4B · XSTest false refusals 2.40% → 4.40% · 98.14% throughput retained

Figure 1: evidence, policy and enforcement are separate modules, so any of them changes without touching request handling.

Adapts the Linux Security Modules separation to LLM serving: a pluggable backend (SAE, transcoder or task-fitted dense probe) emits calibrated per-request evidence, a versioned policy evaluates rules, and a separate gate enforces. Prototyped on Transformers and vLLM. On Qwen3-4B, LMSM-Checkpoint cuts HarmBench attack success from 39.20% to 3.32% while XSTest false refusals rise only from 2.40% to 4.40%, at 98.14% of throughput.

07 of 09

LLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders

mech-interpbackdoorSAEpreprint
Trigger detection (F1) vs. causal control of the language switch, by layer detection F1 ≈ 1causal control (ablation effect) Schematic; detectors and controllers are different feature sets

Figure 1 (schematic): near-perfect detection at many layers, but the features that detect the trigger are not the ones whose ablation removes the behavior.

A controlled backdoor — fixed trigger sequences make 1B and 8B models continue English prompts in French or German — is studied with SAEs across layers and components, contrasting triggered prompts with translation and pretraining controls. SAE features separate triggered prompts from controls with near-perfect F1, but the detecting features do not control the switched behavior: ablating them does not reliably remove the backdoor. Detection and removal need different feature sets.

08 of 09

Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders

mech-interpbackdoorSAEpreprint
clean inputtriggered input clean modelpoisoned model baseline baseline suppressed · weight-modified interaction · mixed

Figure 1 (schematic): backdoor-induced logit shifts are traced to SAE features in four roles; dirty-label and clean-label attacks populate the roles differently.

A 2×2 design (clean vs. poisoned model × clean vs. triggered input) traces backdoor-induced logit shifts to high-contribution SAE features and sorts them into interaction, suppressed, mixed and weight-modified roles. Dirty-label and clean-label attacks populate these roles differently, which is why a defense built for one paradigm misses the other — each existing defense implicitly targets one role.

09 of 09

ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive Alternatives

mech-interpsteeringpreprint
Selective triggerintervene only when needed Refusal-anchoredconstructive redirection Safe completionnot a bare refusal →→ Single inference pass · Llama-3.1 and Qwen2.5

Figure 1 (schematic): a selective gate decides when to intervene; the steering then targets a constructive safe completion.

Frames safety steering as two coupled decisions — when to intervene and how generation proceeds afterwards — and couples a selective trigger with refusal-anchored constructive redirection in one pass. On Llama-3.1 and Qwen2.5 it preserves benign utility while increasing constructive safe-completion behavior, with the largest gains on models that otherwise emit short refusals.

Notes

← all Research Radar issues · gussand · source