Interpretable Spatial-Temporal Forecasting via Additive Neural Decomposition and Knowledge Distillation

Suan Lee, Jinho Kim
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3376-3397, 2026.

Abstract

We challenge the prevailing assumption that interpretability requires sacrificing accuracy in spatial-temporal forecasting. We propose STGNAM (Spatial-Temporal Graph Neural Additive Model), which enforces a strict additive decomposition: $\hat{\mathbf{y}} = f_{\text{node}} + f_{\text{temporal}} + f_{\text{spatial}} + f_{\text{interact}} + \text{bias}$, where each component is independently evaluable and visualizable. Our central finding is that the additive inductive bias acts as an implicit regularizer: on PEMS-BAY, STGNAM *sets a new state of the art* (MAE $1.62$), surpassing all black-box models including TITAN ($1.69$) and its own D2STGNN teacher ($1.87$)—even without knowledge distillation (KD). When combined with KD from a black-box teacher, STGNAM achieves $91$–$104%$ of SOTA across five benchmarks while providing full additive interpretability. KD also produces a secondary benefit: component-level faithfulness increases from $0.53$ to $0.92$ on METR-LA, and skip connection dominance drops from $65%$ to $43%$, indicating that soft-target supervision forces additive components to learn structured, specialized representations. We validate this through direct comparison with post-hoc attribution methods (gradient saliency, SmoothGrad), showing that inherent additive explanations are more faithful and stable than post-hoc alternatives applied to black-box models.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-lee26d, title = {Interpretable Spatial-Temporal Forecasting via Additive Neural Decomposition and Knowledge Distillation}, author = {Lee, Suan and Kim, Jinho}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {3376--3397}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/lee26d/lee26d.pdf}, url = {https://proceedings.mlr.press/v337/lee26d.html}, abstract = {We challenge the prevailing assumption that interpretability requires sacrificing accuracy in spatial-temporal forecasting. We propose STGNAM (Spatial-Temporal Graph Neural Additive Model), which enforces a strict additive decomposition: $\hat{\mathbf{y}} = f_{\text{node}} + f_{\text{temporal}} + f_{\text{spatial}} + f_{\text{interact}} + \text{bias}$, where each component is independently evaluable and visualizable. Our central finding is that the additive inductive bias acts as an implicit regularizer: on PEMS-BAY, STGNAM *sets a new state of the art* (MAE $1.62$), surpassing all black-box models including TITAN ($1.69$) and its own D2STGNN teacher ($1.87$)—even without knowledge distillation (KD). When combined with KD from a black-box teacher, STGNAM achieves $91$–$104%$ of SOTA across five benchmarks while providing full additive interpretability. KD also produces a secondary benefit: component-level faithfulness increases from $0.53$ to $0.92$ on METR-LA, and skip connection dominance drops from $65%$ to $43%$, indicating that soft-target supervision forces additive components to learn structured, specialized representations. We validate this through direct comparison with post-hoc attribution methods (gradient saliency, SmoothGrad), showing that inherent additive explanations are more faithful and stable than post-hoc alternatives applied to black-box models.} }
Endnote
%0 Conference Paper %T Interpretable Spatial-Temporal Forecasting via Additive Neural Decomposition and Knowledge Distillation %A Suan Lee %A Jinho Kim %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-lee26d %I PMLR %P 3376--3397 %U https://proceedings.mlr.press/v337/lee26d.html %V 337 %X We challenge the prevailing assumption that interpretability requires sacrificing accuracy in spatial-temporal forecasting. We propose STGNAM (Spatial-Temporal Graph Neural Additive Model), which enforces a strict additive decomposition: $\hat{\mathbf{y}} = f_{\text{node}} + f_{\text{temporal}} + f_{\text{spatial}} + f_{\text{interact}} + \text{bias}$, where each component is independently evaluable and visualizable. Our central finding is that the additive inductive bias acts as an implicit regularizer: on PEMS-BAY, STGNAM *sets a new state of the art* (MAE $1.62$), surpassing all black-box models including TITAN ($1.69$) and its own D2STGNN teacher ($1.87$)—even without knowledge distillation (KD). When combined with KD from a black-box teacher, STGNAM achieves $91$–$104%$ of SOTA across five benchmarks while providing full additive interpretability. KD also produces a secondary benefit: component-level faithfulness increases from $0.53$ to $0.92$ on METR-LA, and skip connection dominance drops from $65%$ to $43%$, indicating that soft-target supervision forces additive components to learn structured, specialized representations. We validate this through direct comparison with post-hoc attribution methods (gradient saliency, SmoothGrad), showing that inherent additive explanations are more faithful and stable than post-hoc alternatives applied to black-box models.
APA
Lee, S. & Kim, J.. (2026). Interpretable Spatial-Temporal Forecasting via Additive Neural Decomposition and Knowledge Distillation. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:3376-3397 Available from https://proceedings.mlr.press/v337/lee26d.html.

Related Material