DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoML

Xiaoou Ding, Zekai Qian, Siying Chen, Hongbin Hu, Chen Wang, Hongzhi Wang, Jianmin Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:25078-25095, 2026.

Abstract

Data cleaning and automated machine learning (AutoML) are both crucial for reliable learning systems, yet are commonly treated as independent or sequential stages. This separation ignores their strong interaction and leads to inefficient use of limited computational budgets.We propose DMCO, a unified framework that jointly optimizes data cleaning and model construction under a fixed resource budget. DMCO reformulates the traditional two-stage pipeline into a time-sliced process, where data cleaning and AutoML are interleaved and adaptively scheduled. We introduce a gradient-based data cleaning sampling strategy with theoretical guarantees for minimizing gradient estimation variance, and integrates it with loss-driven sampling and progressive AutoML fitting to continuously leverage intermediate data quality improvements.Experiments on six real-world datasets show that DMCO consistently outperforms standalone data cleaning and AutoML baselines on both classification and regression tasks, as measured by F1 score and MSE. Under limited budgets, DMCO achieves up to 82.19% of the performance of full data cleaning with exhaustive AutoML, while remaining robust across different AutoML frameworks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-ding26l, title = {{DMCO}: Budget-Aware Co-Optimization of Data Cleaning and {A}uto{ML}}, author = {Ding, Xiaoou and Qian, Zekai and Chen, Siying and Hu, Hongbin and Wang, Chen and Wang, Hongzhi and Wang, Jianmin}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {25078--25095}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ding26l/ding26l.pdf}, url = {https://proceedings.mlr.press/v306/ding26l.html}, abstract = {Data cleaning and automated machine learning (AutoML) are both crucial for reliable learning systems, yet are commonly treated as independent or sequential stages. This separation ignores their strong interaction and leads to inefficient use of limited computational budgets.We propose DMCO, a unified framework that jointly optimizes data cleaning and model construction under a fixed resource budget. DMCO reformulates the traditional two-stage pipeline into a time-sliced process, where data cleaning and AutoML are interleaved and adaptively scheduled. We introduce a gradient-based data cleaning sampling strategy with theoretical guarantees for minimizing gradient estimation variance, and integrates it with loss-driven sampling and progressive AutoML fitting to continuously leverage intermediate data quality improvements.Experiments on six real-world datasets show that DMCO consistently outperforms standalone data cleaning and AutoML baselines on both classification and regression tasks, as measured by F1 score and MSE. Under limited budgets, DMCO achieves up to 82.19% of the performance of full data cleaning with exhaustive AutoML, while remaining robust across different AutoML frameworks.} }
Endnote
%0 Conference Paper %T DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoML %A Xiaoou Ding %A Zekai Qian %A Siying Chen %A Hongbin Hu %A Chen Wang %A Hongzhi Wang %A Jianmin Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-ding26l %I PMLR %P 25078--25095 %U https://proceedings.mlr.press/v306/ding26l.html %V 306 %X Data cleaning and automated machine learning (AutoML) are both crucial for reliable learning systems, yet are commonly treated as independent or sequential stages. This separation ignores their strong interaction and leads to inefficient use of limited computational budgets.We propose DMCO, a unified framework that jointly optimizes data cleaning and model construction under a fixed resource budget. DMCO reformulates the traditional two-stage pipeline into a time-sliced process, where data cleaning and AutoML are interleaved and adaptively scheduled. We introduce a gradient-based data cleaning sampling strategy with theoretical guarantees for minimizing gradient estimation variance, and integrates it with loss-driven sampling and progressive AutoML fitting to continuously leverage intermediate data quality improvements.Experiments on six real-world datasets show that DMCO consistently outperforms standalone data cleaning and AutoML baselines on both classification and regression tasks, as measured by F1 score and MSE. Under limited budgets, DMCO achieves up to 82.19% of the performance of full data cleaning with exhaustive AutoML, while remaining robust across different AutoML frameworks.
APA
Ding, X., Qian, Z., Chen, S., Hu, H., Wang, C., Wang, H. & Wang, J.. (2026). DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoML. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:25078-25095 Available from https://proceedings.mlr.press/v306/ding26l.html.

Related Material