The field's most widely used metric for validating sparse autoencoder features — cosine similarity between a decoder atom and a ground-truth direction — does not mean the feature is causally active. A rigorous ablation and steering audit reveals that up to 77% of features passing a strict geometric recovery bar are completely causally inert.
SAE evaluation has relied on cosine similarity to ground-truth directions from synthetic toy models, treating high cosine as evidence of successful feature recovery. This work separates the geometric claim ("the decoder direction matches") from the causal claim ("the matched encoder feature fires and has measurable downstream effect"). Zero-ablating features at full layer depth across systematically varied SAE training quality reveals that up to 77% of features with cosine ≥ 0.90 in degraded SAEs — and 9% in well-trained SAEs — are causally inert: the decoder atom aligns perfectly while the encoder never fires when the feature is present. The study also reproduces the Elhage et al. (2022) superposition phase diagram and identifies a convergence artifact at high sparsity and a previously undescribed diffuse sharing regime at extreme overcompleteness that inflate cosine metrics without producing causally active features.