Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability

Amir Asiaee
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:210-227, 2026.

Abstract

Mechanistic interpretability often evaluates explanations by intervening on a model: swapping hidden states, patching activations, ablating components, or comparing a compressed model to the original one. These experiments are usually summarized by a point estimate, even though the evaluation may be monitored while it runs or adapted toward suspected failures. This makes it hard to tell whether a reported fidelity or patching effect is a stable causal claim or a consequence of finite sampling and evaluation choices. We introduce Certified Interventional Fidelity ({CIF}), a statistical layer for interventional interpretability evaluations. {CIF} first writes the quantity being reported as a causal estimand: an expectation of a bounded score over a stated input distribution and a stated intervention distribution. It then provides confidence intervals and anytime-valid confidence sequences for this estimand, including under adaptive intervention sampling via bounded mixture importance weighting. We instantiate {CIF} with {Hoeffding}-style sequences and variance-adaptive betting sequences, the latter reducing certification cost by $10$–$30\times$ in our experiments. On {MNIST} abstractions and {GPT-2} Small {IOI} circuits, {CIF} certifies high-fidelity claims, shows when apparent method differences are not statistically supported, and makes sensitivity to the intervention distribution explicit.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-asiaee26d, title = {Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability}, author = {Asiaee, Amir}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {210--227}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/asiaee26d/asiaee26d.pdf}, url = {https://proceedings.mlr.press/v337/asiaee26d.html}, abstract = {Mechanistic interpretability often evaluates explanations by intervening on a model: swapping hidden states, patching activations, ablating components, or comparing a compressed model to the original one. These experiments are usually summarized by a point estimate, even though the evaluation may be monitored while it runs or adapted toward suspected failures. This makes it hard to tell whether a reported fidelity or patching effect is a stable causal claim or a consequence of finite sampling and evaluation choices. We introduce Certified Interventional Fidelity ({CIF}), a statistical layer for interventional interpretability evaluations. {CIF} first writes the quantity being reported as a causal estimand: an expectation of a bounded score over a stated input distribution and a stated intervention distribution. It then provides confidence intervals and anytime-valid confidence sequences for this estimand, including under adaptive intervention sampling via bounded mixture importance weighting. We instantiate {CIF} with {Hoeffding}-style sequences and variance-adaptive betting sequences, the latter reducing certification cost by $10$–$30\times$ in our experiments. On {MNIST} abstractions and {GPT-2} Small {IOI} circuits, {CIF} certifies high-fidelity claims, shows when apparent method differences are not statistically supported, and makes sensitivity to the intervention distribution explicit.} }
Endnote
%0 Conference Paper %T Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability %A Amir Asiaee %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-asiaee26d %I PMLR %P 210--227 %U https://proceedings.mlr.press/v337/asiaee26d.html %V 337 %X Mechanistic interpretability often evaluates explanations by intervening on a model: swapping hidden states, patching activations, ablating components, or comparing a compressed model to the original one. These experiments are usually summarized by a point estimate, even though the evaluation may be monitored while it runs or adapted toward suspected failures. This makes it hard to tell whether a reported fidelity or patching effect is a stable causal claim or a consequence of finite sampling and evaluation choices. We introduce Certified Interventional Fidelity ({CIF}), a statistical layer for interventional interpretability evaluations. {CIF} first writes the quantity being reported as a causal estimand: an expectation of a bounded score over a stated input distribution and a stated intervention distribution. It then provides confidence intervals and anytime-valid confidence sequences for this estimand, including under adaptive intervention sampling via bounded mixture importance weighting. We instantiate {CIF} with {Hoeffding}-style sequences and variance-adaptive betting sequences, the latter reducing certification cost by $10$–$30\times$ in our experiments. On {MNIST} abstractions and {GPT-2} Small {IOI} circuits, {CIF} certifies high-fidelity claims, shows when apparent method differences are not statistically supported, and makes sensitivity to the intervention distribution explicit.
APA
Asiaee, A.. (2026). Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:210-227 Available from https://proceedings.mlr.press/v337/asiaee26d.html.

Related Material