Auditing Feature Importance Rankings in Nonstationary Data Streams

Bahareh Daneshvar, Conor Fahy, Shengxiang Yang
Proceedings of the Fourth UK AI Conference 2026, PMLR 348:42-51, 2026.

Abstract

Feature importance explanations are often evaluated on fixed tabular data, although many deployed models operate on nonstationary streams in which the data distribution and relevant signal may change over time. In such settings, a ranking can appear stable while failing to respond to drift, or change for reasons unrelated to the label signal learned by the model. This paper proposes a protocol for evaluating feature importance rankings in data streams. The protocol combines a sanity gate based on label randomization, deletion tests for faithfulness, keep-top-k sufficiency, temporal rank stability, negative controls, and Drift–Explanation Alignment (DEA), a diagnostic that relates distributional drift to changes in feature rankings. We evaluate the protocol on controlled synthetic streams representing stable, abrupt, gradual, and correlated proxy conditions, as well as selected real datasets from electricity markets, insect monitoring, gas sensing, and industrial fault detection. Controls that are independent of the trained model fail the sanity gate, even when they are perfectly stable or responsive to drift. Rankings derived from trained models recover known signal features in controlled streams but show mixed reliability on real data. These findings support interpreting DEA and temporal stability only after establishing model dependence.

Cite this Paper


BibTeX
@InProceedings{pmlr-v348-daneshvar26a, title = {Auditing Feature Importance Rankings in Nonstationary Data Streams}, author = {Daneshvar, Bahareh and Fahy, Conor and Yang, Shengxiang}, booktitle = {Proceedings of the Fourth UK AI Conference 2026}, pages = {42--51}, year = {2026}, editor = {Benford, Alistair and Büyükateş, Baturalp and Cabrera, Christian and Kiden, Sarah and Salili-James, Arianna and Zakka, Vincent and Zhou, Feng}, volume = {348}, series = {Proceedings of Machine Learning Research}, month = {29--30 Sep}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v348/main/assets/daneshvar26a/daneshvar26a.pdf}, url = {https://proceedings.mlr.press/v348/daneshvar26a.html}, abstract = {Feature importance explanations are often evaluated on fixed tabular data, although many deployed models operate on nonstationary streams in which the data distribution and relevant signal may change over time. In such settings, a ranking can appear stable while failing to respond to drift, or change for reasons unrelated to the label signal learned by the model. This paper proposes a protocol for evaluating feature importance rankings in data streams. The protocol combines a sanity gate based on label randomization, deletion tests for faithfulness, keep-top-k sufficiency, temporal rank stability, negative controls, and Drift–Explanation Alignment (DEA), a diagnostic that relates distributional drift to changes in feature rankings. We evaluate the protocol on controlled synthetic streams representing stable, abrupt, gradual, and correlated proxy conditions, as well as selected real datasets from electricity markets, insect monitoring, gas sensing, and industrial fault detection. Controls that are independent of the trained model fail the sanity gate, even when they are perfectly stable or responsive to drift. Rankings derived from trained models recover known signal features in controlled streams but show mixed reliability on real data. These findings support interpreting DEA and temporal stability only after establishing model dependence.} }
Endnote
%0 Conference Paper %T Auditing Feature Importance Rankings in Nonstationary Data Streams %A Bahareh Daneshvar %A Conor Fahy %A Shengxiang Yang %B Proceedings of the Fourth UK AI Conference 2026 %C Proceedings of Machine Learning Research %D 2026 %E Alistair Benford %E Baturalp Büyükateş %E Christian Cabrera %E Sarah Kiden %E Arianna Salili-James %E Vincent Zakka %E Feng Zhou %F pmlr-v348-daneshvar26a %I PMLR %P 42--51 %U https://proceedings.mlr.press/v348/daneshvar26a.html %V 348 %X Feature importance explanations are often evaluated on fixed tabular data, although many deployed models operate on nonstationary streams in which the data distribution and relevant signal may change over time. In such settings, a ranking can appear stable while failing to respond to drift, or change for reasons unrelated to the label signal learned by the model. This paper proposes a protocol for evaluating feature importance rankings in data streams. The protocol combines a sanity gate based on label randomization, deletion tests for faithfulness, keep-top-k sufficiency, temporal rank stability, negative controls, and Drift–Explanation Alignment (DEA), a diagnostic that relates distributional drift to changes in feature rankings. We evaluate the protocol on controlled synthetic streams representing stable, abrupt, gradual, and correlated proxy conditions, as well as selected real datasets from electricity markets, insect monitoring, gas sensing, and industrial fault detection. Controls that are independent of the trained model fail the sanity gate, even when they are perfectly stable or responsive to drift. Rankings derived from trained models recover known signal features in controlled streams but show mixed reliability on real data. These findings support interpreting DEA and temporal stability only after establishing model dependence.
APA
Daneshvar, B., Fahy, C. & Yang, S.. (2026). Auditing Feature Importance Rankings in Nonstationary Data Streams. Proceedings of the Fourth UK AI Conference 2026, in Proceedings of Machine Learning Research 348:42-51 Available from https://proceedings.mlr.press/v348/daneshvar26a.html.

Related Material