📡 Research Radar · Daily

Mech Interp · AI Security · Text Diffusion LMs

August 9, 2026  ·  Window: Aug 7–9, 2026 + high-priority backlog
Sources: arXiv (cs.CL / cs.LG / cs.CR / cs.AI), IEEE Access, OpenReview
1 peer-reviewed 9 preprints 0 forum/blog
01 mech-interp AI security IEEE Access 2026

Detecting Safety Training Modification in Language Models via Activation Analysis

Can you fingerprint how a model's safety training was modified — purely from its activation geometry, no behavioral queries needed? AMS (Activation-based Model Scanner) says yes. Safety fine-tuning creates measurable separation between harmful and benign content classes in activation space; abliteration and uncensored fine-tuning each collapse or rotate that structure in characteristic, detectable ways. The tool is fully black-box to weights, operating entirely on intermediate-layer activation patterns at inference time.

Safety-Trained Model Abliterated Model PC1 → PC2 ↑ Benign Harmful large gap Collapsed AMS: 71% LOOCV · 14 configs · 4 families
Figure 1 — AMS probes the geometric structure of harmful/benign concept clusters in activation space. Safety-trained models (left) show clear separation; abliteration collapses the clusters into a mixed region (right), a signal AMS uses to classify modification type at 71% leave-one-out accuracy.

AMS extracts intermediate-layer residual-stream activations for paired harmful/benign prompts and computes linear probe separability, cosine distances, and angular distribution statistics across the activation manifold. Validated across 14 model configurations spanning Llama, Gemma, Qwen, and Mistral in four safety-modification categories (instruction-tuned, base, abliterated, uncensored fine-tunes), AMS achieves 71% leave-one-out cross-validation accuracy in classifying which modification the model underwent — without any behavioral queries or weight access. That different modification types leave geometrically distinct footprints opens the door to passive, inference-time model auditing, complementing behavioral red-teaming with a structural signal.

Items 04 – 10 · Condensed
05 AI security preprint · Aug 2026

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

a₁ a₂ a₃⚠ a₄ ✗ z₁ z₂ z₃ z₄ ⛔ BLOCK risk →
Recurrent latent state z accumulates prefix-risk; trajectory-level block at a₄.

Addresses the critical blind spot in per-action guardrails: individually benign-looking actions can gradually drift an LLM agent toward hazardous states in long-horizon tasks. DreamGuard maintains a compact recurrent latent encoder over the action trajectory, predicts future latent states, and derives both immediate-hazard and prefix-risk scores before each tool invocation — empirically catching multi-step unsafe trajectories that point-in-time guards consistently miss.

09 mech-interp preprint · Jul 2026

Stress Testing Concept Erasure with Large Language Model Agents

Proposer gen tests Critique challenge Verify confirm External knowledge
Iterative propose → challenge → verify loop with external knowledge retrieval.

STACE autonomously stress-tests concept-erased models via a three-agent loop: Proposer generates natural-language probes grounded by external knowledge, Critique challenges weak tests, Verifier confirms whether erased concepts remain recoverable. Exposes failure modes invisible to static benchmark sets, particularly under diverse rephrasing and cross-concept knowledge reconstruction.

7 entries removed on 2026-09-10 as repeats of earlier reports: 2607.15893 (first covered 2026-07-21), 2606.20560 (first covered 2026-07-01), 2607.07903 (first covered 2026-07-12), 2605.10971 (first covered 2026-07-01), 2606.06333 (first covered 2026-07-02), 2606.04027 (first covered 2026-07-02), 2605.19262 (first covered 2026-07-03).

← all Research Radar issues · gussand · source