Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability
Activation-patching and ablation results in mech-interp papers are routinely reported as point estimates — even when evaluations were monitored and adapted mid-run, which silently inflates false-positive rates with no principled bound. This UAI 2026 paper gives the field a rigorous statistical layer for the first time.
CIF writes the evaluated quantity as a causal estimand — an expectation of a bounded score over a stated input distribution and a stated intervention distribution. Confidence intervals and anytime-valid confidence sequences follow via bounded mixture importance weighting. Because the sequences remain valid at any sample size, practitioners can halt early when the effect is conclusive or extend when it is ambiguous, without inflating type-I error. The framework explicitly handles adaptive intervention sampling — the dominant evaluation pattern in practice — covering activation patching, state swapping, component ablation, and compressed-model comparison under a unified formulation.