Research Radar
Daily · September 8, 2026
2 peer-reviewed · 5 preprints · 0 forum/blog
Pretraining Safety · AI Security · Applied Mech Interp

02
EMNLP 2026 mech-interp refusal circuits

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

Hoang Cuong Nguyen, Mark Dras, Usman Naseem — EMNLP 2026 Main Conference (peer-reviewed)

Three post-training recipes — SFT, reasoning-augmented SFT, and ORPO — produce models that refuse harmful requests at similar rates, yet install fundamentally different internal circuits to do it. This paper measures those differences mechanistically across three model families and finds that no current method achieves all three properties a durably safe model needs at once.

SFT Reasoning-SFT ORPO Concentrated few heads Easily ablated Fragile to steering Distinct circuit all 3 models More distributed Capability cost Intermediate fragility Arch-dependent Less cap. cost Structure Robustness Capability No single method achieves all three: non-fragile · no capability cost · correctable
Figure 2: Refusal circuit comparison across SFT, reasoning-augmented SFT, and ORPO on Llama-3.1-8B, Gemma-2-9B, Qwen3-8B. No method simultaneously achieves non-concentrated, capability-preserving, and steerable refusal.

Using attention head attribution and activation patching, the study finds that SFT concentrates refusal in a small, easily ablated set of heads; reasoning-augmented training distributes refusal more broadly and produces a consistently distinct computational signature across all three model families; ORPO sits between the two on fragility but sacrifices less capability. Architecture has an independent effect: the same training method installs differently steerable circuits depending on the base model. The trilemma — non-concentrated refusal, no capability regression, correctable refusal — holds across all combinations tested, challenging the assumption that better post-training data alone can close safety gaps visible only at the circuit level.



Items 4–10 · Also notable

05
mech-interp preprint

Locating and Steering Refusal Beyond Attention

Preethi Carmel Bosco, Gopalakrishnan Srinivasan — arXiv, September 4, 2026
Transformer attention heads ↑ refusal dir. rigid rotation R ∈ SO(d) SSM recurrent update ↑ refusal dir. Probe trained on transformer ✓ flags SSM harmful inputs · ablation removes SSM refusal
Figure 5: A rigid rotation aligns transformer and SSM residual streams; the shared refusal direction persists across fundamentally different token-mixing architectures, enabling cross-model probe transfer.

State-space models share a refusal direction with transformers despite using recurrent rather than attention-based token mixing. A rigid rotation aligns representation spaces; a harm probe trained on a transformer then flags SSM harmful inputs without re-training, and ablating the aligned direction removes SSM refusal — showing the safety representation is architecture-agnostic.


06
AI security open-weight preprint

Uncensored Open-weight Models: Redistribution as the Persistence Layer

Juliette Garcia, Hailey May et al. (10a Labs) — arXiv, September 4, 2026
Uncensored models on HuggingFace (Jan 2024 – Mar 2026) count Q1'24 Q1'25 Q1'26 3,471 original models total × 2.4 avg repacks = 8,164 redistributions 25% of 1,643 apps explicitly malicious
Figure 6: Rapid growth of uncensored open-weight models 2024–2026. 3,471 original models, each repackaged ~2.4 times; 3 actors account for 52% of all 8,164 redistributions. Once quantized and mirrored, models persist across upstream removal.

Between January 2024 and March 2026, 3,471 original uncensored models appeared on HuggingFace, each repackaged 2.4 times on average; three actors account for 52% of all 8,164 compressed redistributions. Redistribution across Ollama and alternative registries creates a persistence layer that survives upstream removal: of 1,643 identified GitHub integrations, 25% are classified as explicitly malicious.


07
mech-interp benchmark preprint

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

arXiv, September 2, 2026
Observer (interp method) Estimation accuracy pairwise: high ↑ Action-choice quality not always better ↓ accuracy– control gap task-dependent ObserverBench: evaluation framework
Figure 7: ObserverBench decouples estimation accuracy from action-choice quality. On circuit-intervention tasks in GPT-2-small and Qwen2.5-7B, pairwise observers achieve high prediction accuracy but do not always choose lower-loss actions — the accuracy–control gap is substantial and task-dependent.

Mechanistic interpretability methods increasingly guide real interventions, but an estimate accurate on average can choose poor actions in context. ObserverBench standardizes evaluation: each task fixes model, information boundary, allowed actions, decision rule, and held-out cases, reporting estimation accuracy separately from action-choice loss. On circuit-intervention tasks, pairwise observers predict unseen effects better but do not consistently select lower-loss actions, exposing a systematic accuracy–control gap.


08
ECCV 2026 concept erasure SAE

EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders

HiDream-ai team — ECCV 2026 (peer-reviewed)
1. Decompose Partitioned Conv. SAE → sparse monosemantic feats. 2. Attribute contrastive pairs → concept kernels 3. Erase timestep-resolved spatial masks confine erasure HunyuanVideo · CogVideoX-5b — celebrity identity & nudity erasure
Figure 8: EraseSAE pipeline — Partitioned Convolutional SAE decomposes spatiotemporal activations into monosemantic features; contrastive attribution isolates concept-specific kernels; timestep-resolved masks confine erasure to active regions, leaving adjacent concepts intact.

Coarse concept-erasure methods degrade adjacent content because they do not operate at the level of individual monosemantic features. EraseSAE decomposes DiT activations with a Partitioned Convolutional SAE, uses contrastive paired prompts to isolate concept-specific feature kernels, and applies timestep-resolved spatial masks at inference — achieving precise celebrity-identity and nudity erasure on HunyuanVideo and CogVideoX-5b with minimal quality regression.


09
AI security preprint

The Implications of Linguistic Illegibility for LLM Security

James Mickens (Harvard) — arXiv, September 2, 2026
Internal computation math over activation spaces not directly in language true intent here lossy translation Linguistic externalization CoT · self-critique · probes read only this layer may miss true intent ✗ Linguistic illegibility → CoT monitoring, constitutional critique, activation probes all unreliable
Figure 9: Linguistic illegibility — internal computation occurs in activation space, not in natural language; the translation to linguistic output is lossy, so security mechanisms that read only linguistic artifacts may fail to detect internally computed intents.

"Linguistic illegibility" describes scenarios where an LLM's externalized language fails to represent its internal computation. The paper argues this is unavoidable when internal math over activation spaces is translated lossily to language, and catalogs which current LLM security mechanisms — chain-of-thought monitoring, constitutional self-critique, linguistically-defined activation probing — are therefore unreliable by construction.


10
alignment pretraining safety preprint

Representational Alignment Yields Generalizable Safety in Language Models

Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu — arXiv, September 3, 2026

3 entries removed on 2026-09-10 as repeats of earlier reports: 2608.11025 (first covered 2026-08-20), 2609.00051 (first covered 2026-09-03), 2609.02293 (first covered 2026-09-05).

Standard RLHF aligns output responses latent space: unaligned fails on adversarial rephrasings RSO aligns latent space to moral prototypes generalizes across rephrasings Human moral prototypes typicality ratings 23 LLMs tested — weak baseline moral typicality preservation; RSO improves adversarial generalization
Figure 10: Standard RLHF aligns outputs but leaves latent moral representations unaligned; RSO directly targets the latent space geometry using human moral typicality ratings, improving generalization to adversarially-recast harmful inputs across 23 LLMs.

Standard RLHF aligns model outputs but not the underlying latent representation of harm categories, leaving models vulnerable when harmful intent is recast in adversarial forms. Representational Similarity Optimization (RSO) directly aligns LLM latent geometry with human moral judgment prototypes; evaluation across 23 LLMs shows that baseline moral typicality is weakly preserved, and RSO training substantially improves generalization to adversarial harm rephrasings.

Notes

← all Research Radar issues · gussand · source