Filter, Augment, Forecast: Online Data Selection for Robust Time Series Forecasting

Ege Onur Taga, Halil Alperen Gozeten, Kutay Tire, Rahul Dalvi, Reinhard Heckel, Samet Oymak
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:1081-1089, 2026.

Abstract

Data curation pipelines play a central role in training deep learning architectures, with their impact in time series forecasting still relatively underexplored. In this work, we propose Filter, Augment, Forecast (FAF): an online data curation strategy based on (1) data selection to filter out low-quality (e.g., noisy) examples and (2) augmentation of the remaining high-quality data. We use reference model-based filtering inspired by the reducible holdout loss selection (RHO-LOSS) from the language modeling literature. We identify limitations of RHO-LOSS under domain shifts common in time series and introduce the adaptive RHO method (AdaRho), which improves performance by updating the reference model during training. Using random matrix theory, we provide a statistical analysis that characterizes the role of the reference model, sample size, and noise statistics in data selection. FAF consistently improves forecasting accuracy across diverse architectures without modifying them, achieving state-of-the-art results.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-taga26a, title = { Filter, Augment, Forecast: Online Data Selection for Robust Time Series Forecasting }, author = {Taga, Ege Onur and Gozeten, Halil Alperen and Tire, Kutay and Dalvi, Rahul and Heckel, Reinhard and Oymak, Samet}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {1081--1089}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/taga26a/taga26a.pdf}, url = {https://proceedings.mlr.press/v300/taga26a.html}, abstract = { Data curation pipelines play a central role in training deep learning architectures, with their impact in time series forecasting still relatively underexplored. In this work, we propose Filter, Augment, Forecast (FAF): an online data curation strategy based on (1) data selection to filter out low-quality (e.g., noisy) examples and (2) augmentation of the remaining high-quality data. We use reference model-based filtering inspired by the reducible holdout loss selection (RHO-LOSS) from the language modeling literature. We identify limitations of RHO-LOSS under domain shifts common in time series and introduce the adaptive RHO method (AdaRho), which improves performance by updating the reference model during training. Using random matrix theory, we provide a statistical analysis that characterizes the role of the reference model, sample size, and noise statistics in data selection. FAF consistently improves forecasting accuracy across diverse architectures without modifying them, achieving state-of-the-art results. } }
Endnote
%0 Conference Paper %T Filter, Augment, Forecast: Online Data Selection for Robust Time Series Forecasting %A Ege Onur Taga %A Halil Alperen Gozeten %A Kutay Tire %A Rahul Dalvi %A Reinhard Heckel %A Samet Oymak %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-taga26a %I PMLR %P 1081--1089 %U https://proceedings.mlr.press/v300/taga26a.html %V 300 %X Data curation pipelines play a central role in training deep learning architectures, with their impact in time series forecasting still relatively underexplored. In this work, we propose Filter, Augment, Forecast (FAF): an online data curation strategy based on (1) data selection to filter out low-quality (e.g., noisy) examples and (2) augmentation of the remaining high-quality data. We use reference model-based filtering inspired by the reducible holdout loss selection (RHO-LOSS) from the language modeling literature. We identify limitations of RHO-LOSS under domain shifts common in time series and introduce the adaptive RHO method (AdaRho), which improves performance by updating the reference model during training. Using random matrix theory, we provide a statistical analysis that characterizes the role of the reference model, sample size, and noise statistics in data selection. FAF consistently improves forecasting accuracy across diverse architectures without modifying them, achieving state-of-the-art results.
APA
Taga, E.O., Gozeten, H.A., Tire, K., Dalvi, R., Heckel, R. & Oymak, S.. (2026). Filter, Augment, Forecast: Online Data Selection for Robust Time Series Forecasting . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:1081-1089 Available from https://proceedings.mlr.press/v300/taga26a.html.

Related Material