Emergent misalignment — where a model trained on one task (coding) suddenly starts giving harmful advice — turns out to be mechanistically legible: specific SAE persona features amplified during fine-tuning cause it, and steering just those features achieves a higher misalignment rate than fine-tuning the whole model.
Applying SAE-based model diffing across 4 open-weight models, the authors identify features that are systematically up- or down-regulated by misalignment-inducing fine-tuning. Amplified features correspond to "jailbreak persona," sarcasm, deception, and manipulation; safety-relevant and assistant-identity features are suppressed. Steering individual identified features via activation addition achieves a 62% misalignment rate, compared to 35% from the fine-tuning procedure itself. The result implies that emergent misalignment is not a holistic distributional shift but a narrow mechanistic perturbation — suggesting that monitoring or correcting a small set of persona features could detect or reverse it.