Hoang Cuong Nguyen, Mark Dras, Usman Naseem — EMNLP 2026 Main Conference (peer-reviewed)
Three post-training recipes — SFT, reasoning-augmented SFT, and ORPO — produce models that refuse harmful requests at similar rates, yet install fundamentally different internal circuits to do it. This paper measures those differences mechanistically across three model families and finds that no current method achieves all three properties a durably safe model needs at once.
Figure 2: Refusal circuit comparison across SFT, reasoning-augmented SFT, and ORPO on Llama-3.1-8B, Gemma-2-9B, Qwen3-8B. No method simultaneously achieves non-concentrated, capability-preserving, and steerable refusal.
Using attention head attribution and activation patching, the study finds that SFT concentrates refusal in a small, easily ablated set of heads; reasoning-augmented training distributes refusal more broadly and produces a consistently distinct computational signature across all three model families; ORPO sits between the two on fragility but sacrifices less capability. Architecture has an independent effect: the same training method installs differently steerable circuits depending on the base model. The trilemma — non-concentrated refusal, no capability regression, correctable refusal — holds across all combinations tested, challenging the assumption that better post-training data alone can close safety gaps visible only at the circuit level.
Preethi Carmel Bosco, Gopalakrishnan Srinivasan — arXiv, September 4, 2026
Figure 5: A rigid rotation aligns transformer and SSM residual streams; the shared refusal direction persists across fundamentally different token-mixing architectures, enabling cross-model probe transfer.
State-space models share a refusal direction with transformers despite using recurrent rather than attention-based token mixing. A rigid rotation aligns representation spaces; a harm probe trained on a transformer then flags SSM harmful inputs without re-training, and ablating the aligned direction removes SSM refusal — showing the safety representation is architecture-agnostic.
Juliette Garcia, Hailey May et al. (10a Labs) — arXiv, September 4, 2026
Figure 6: Rapid growth of uncensored open-weight models 2024–2026. 3,471 original models, each repackaged ~2.4 times; 3 actors account for 52% of all 8,164 redistributions. Once quantized and mirrored, models persist across upstream removal.
Between January 2024 and March 2026, 3,471 original uncensored models appeared on HuggingFace, each repackaged 2.4 times on average; three actors account for 52% of all 8,164 compressed redistributions. Redistribution across Ollama and alternative registries creates a persistence layer that survives upstream removal: of 1,643 identified GitHub integrations, 25% are classified as explicitly malicious.
Figure 7: ObserverBench decouples estimation accuracy from action-choice quality. On circuit-intervention tasks in GPT-2-small and Qwen2.5-7B, pairwise observers achieve high prediction accuracy but do not always choose lower-loss actions — the accuracy–control gap is substantial and task-dependent.
Mechanistic interpretability methods increasingly guide real interventions, but an estimate accurate on average can choose poor actions in context. ObserverBench standardizes evaluation: each task fixes model, information boundary, allowed actions, decision rule, and held-out cases, reporting estimation accuracy separately from action-choice loss. On circuit-intervention tasks, pairwise observers predict unseen effects better but do not consistently select lower-loss actions, exposing a systematic accuracy–control gap.
Figure 8: EraseSAE pipeline — Partitioned Convolutional SAE decomposes spatiotemporal activations into monosemantic features; contrastive attribution isolates concept-specific kernels; timestep-resolved masks confine erasure to active regions, leaving adjacent concepts intact.
Coarse concept-erasure methods degrade adjacent content because they do not operate at the level of individual monosemantic features. EraseSAE decomposes DiT activations with a Partitioned Convolutional SAE, uses contrastive paired prompts to isolate concept-specific feature kernels, and applies timestep-resolved spatial masks at inference — achieving precise celebrity-identity and nudity erasure on HunyuanVideo and CogVideoX-5b with minimal quality regression.
James Mickens (Harvard) — arXiv, September 2, 2026
Figure 9: Linguistic illegibility — internal computation occurs in activation space, not in natural language; the translation to linguistic output is lossy, so security mechanisms that read only linguistic artifacts may fail to detect internally computed intents.
"Linguistic illegibility" describes scenarios where an LLM's externalized language fails to represent its internal computation. The paper argues this is unavoidable when internal math over activation spaces is translated lossily to language, and catalogs which current LLM security mechanisms — chain-of-thought monitoring, constitutional self-critique, linguistically-defined activation probing — are therefore unreliable by construction.
Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu — arXiv, September 3, 2026
3 entries removed on 2026-09-10 as repeats of earlier reports: 2608.11025 (first covered 2026-08-20), 2609.00051 (first covered 2026-09-03), 2609.02293 (first covered 2026-09-05).
Figure 10: Standard RLHF aligns outputs but leaves latent moral representations unaligned; RSO directly targets the latent space geometry using human moral typicality ratings, improving generalization to adversarially-recast harmful inputs across 23 LLMs.
Standard RLHF aligns model outputs but not the underlying latent representation of harm categories, leaving models vulnerable when harmful intent is recast in adversarial forms. Representational Similarity Optimization (RSO) directly aligns LLM latent geometry with human moral judgment prototypes; evaluation across 23 LLMs shows that baseline moral typicality is weakly preserved, and RSO training substantially improves generalization to adversarial harm rephrasings.
Notes
Window: September 6–8, 2026 primary; EMNLP 2026 Main/Findings papers submitted September 1–5 included as newly posted peer-reviewed results not in the September 6 radar.
Deep Ignorance (2508.06601, ICLR 2026) was the #1 item in the September 6 report and is not re-listed here.
Items 1, 2, 3, 4, and 8 are flagged for the W37 weekly roundup.
No LessWrong/Alignment Forum posts of sufficient relevance found for this window.