Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

Wanjin Feng, Yuan Yuan, Jingtao Ding, Yong Li
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:30490-30509, 2026.

Abstract

In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, standard evaluations rely on aggregate metrics (e.g., MSE) that conflate model capability with the intrinsic difficulty of the evaluated instances. To address this, we propose a diagnostic framework anchored in Spectral Coherence Predictability (SCP), which provides an efficient $\mathcal{O}(N\log N)$ per-instance difficulty reference and yields a corresponding linear MSE lower bound. Complementing this, we introduce the Linear Utilization Ratio (LUR) to quantify how effectively models exploit linearly predictable structures across frequencies. Experiments on synthetic and real-world benchmarks show that SCP aligns strongly with realized forecasting errors across diverse state-of-the-art forecasters. Using this lens, we uncover “predictability drift,” revealing that task difficulty is not static but fluctuates significantly over time and variables. Furthermore, stratified evaluation exposes complementary architectural strengths across distinct frequency bands and difficulty regimes. Overall, we advocate moving beyond leaderboard-style ranking toward a more insightful, predictability-aware evaluation that fosters fairer model comparisons and a deeper understanding of model behavior. Code and data are available at https://github.com/WanjinVon/TS_Predictability.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-feng26u, title = {Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting}, author = {Feng, Wanjin and Yuan, Yuan and Ding, Jingtao and Li, Yong}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {30490--30509}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/feng26u/feng26u.pdf}, url = {https://proceedings.mlr.press/v306/feng26u.html}, abstract = {In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, standard evaluations rely on aggregate metrics (e.g., MSE) that conflate model capability with the intrinsic difficulty of the evaluated instances. To address this, we propose a diagnostic framework anchored in Spectral Coherence Predictability (SCP), which provides an efficient $\mathcal{O}(N\log N)$ per-instance difficulty reference and yields a corresponding linear MSE lower bound. Complementing this, we introduce the Linear Utilization Ratio (LUR) to quantify how effectively models exploit linearly predictable structures across frequencies. Experiments on synthetic and real-world benchmarks show that SCP aligns strongly with realized forecasting errors across diverse state-of-the-art forecasters. Using this lens, we uncover “predictability drift,” revealing that task difficulty is not static but fluctuates significantly over time and variables. Furthermore, stratified evaluation exposes complementary architectural strengths across distinct frequency bands and difficulty regimes. Overall, we advocate moving beyond leaderboard-style ranking toward a more insightful, predictability-aware evaluation that fosters fairer model comparisons and a deeper understanding of model behavior. Code and data are available at https://github.com/WanjinVon/TS_Predictability.} }
Endnote
%0 Conference Paper %T Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting %A Wanjin Feng %A Yuan Yuan %A Jingtao Ding %A Yong Li %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-feng26u %I PMLR %P 30490--30509 %U https://proceedings.mlr.press/v306/feng26u.html %V 306 %X In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, standard evaluations rely on aggregate metrics (e.g., MSE) that conflate model capability with the intrinsic difficulty of the evaluated instances. To address this, we propose a diagnostic framework anchored in Spectral Coherence Predictability (SCP), which provides an efficient $\mathcal{O}(N\log N)$ per-instance difficulty reference and yields a corresponding linear MSE lower bound. Complementing this, we introduce the Linear Utilization Ratio (LUR) to quantify how effectively models exploit linearly predictable structures across frequencies. Experiments on synthetic and real-world benchmarks show that SCP aligns strongly with realized forecasting errors across diverse state-of-the-art forecasters. Using this lens, we uncover “predictability drift,” revealing that task difficulty is not static but fluctuates significantly over time and variables. Furthermore, stratified evaluation exposes complementary architectural strengths across distinct frequency bands and difficulty regimes. Overall, we advocate moving beyond leaderboard-style ranking toward a more insightful, predictability-aware evaluation that fosters fairer model comparisons and a deeper understanding of model behavior. Code and data are available at https://github.com/WanjinVon/TS_Predictability.
APA
Feng, W., Yuan, Y., Ding, J. & Li, Y.. (2026). Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:30490-30509 Available from https://proceedings.mlr.press/v306/feng26u.html.

Related Material