Can you fingerprint how a model's safety training was modified — purely from its activation geometry, no behavioral queries needed? AMS (Activation-based Model Scanner) says yes. Safety fine-tuning creates measurable separation between harmful and benign content classes in activation space; abliteration and uncensored fine-tuning each collapse or rotate that structure in characteristic, detectable ways. The tool is fully black-box to weights, operating entirely on intermediate-layer activation patterns at inference time.
Figure 1 — AMS probes the geometric structure of harmful/benign concept clusters in activation space. Safety-trained models (left) show clear separation; abliteration collapses the clusters into a mixed region (right), a signal AMS uses to classify modification type at 71% leave-one-out accuracy.
AMS extracts intermediate-layer residual-stream activations for paired harmful/benign prompts and computes linear probe separability, cosine distances, and angular distribution statistics across the activation manifold. Validated across 14 model configurations spanning Llama, Gemma, Qwen, and Mistral in four safety-modification categories (instruction-tuned, base, abliterated, uncensored fine-tunes), AMS achieves 71% leave-one-out cross-validation accuracy in classifying which modification the model underwent — without any behavioral queries or weight access. That different modification types leave geometrically distinct footprints opens the door to passive, inference-time model auditing, complementing behavioral red-teaming with a structural signal.
Wenhao Lin, Chenyu Yu et al. — arXiv:2608.05695, August 6, 2026
Recurrent latent state z accumulates prefix-risk; trajectory-level block at a₄.
Addresses the critical blind spot in per-action guardrails: individually benign-looking actions can gradually drift an LLM agent toward hazardous states in long-horizon tasks. DreamGuard maintains a compact recurrent latent encoder over the action trajectory, predicts future latent states, and derives both immediate-hazard and prefix-risk scores before each tool invocation — empirically catching multi-step unsafe trajectories that point-in-time guards consistently miss.