Undersmoothing Black-Box Models for Functional Estimation

Yue Yu, Debarghya Mukherjee, Moulinath Banerjee, Yaacov Ritov
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:2296-2304, 2026.

Abstract

We study functional estimation using black-box models through a model-agnostic undersmoothing framework. The proposed procedure \texttt{Rep} operates by augmenting the original dataset through replicating a proportion of samples multiple times, and subsequently applying the black-box algorithm to the augmented dataset. This construction automatically induces undersmoothing and reduces the functional estimation error. We provide several empirical demonstrations (including neural network based learners) showing that compared to the plug-in estimator, the proposed algorithm \texttt{Rep} improves the estimation accuracy of functional estimation without requiring explicit expressions for the associated influence functions. Furthermore, we develop a theoretical analysis in two representative settings, the Nadaraya–Watson estimator and the random feature model, establishing that replication provides explicit prescriptions for the replication proportion and number of copies, and yields optimal convergence rates for functional estimation. In the classical nonparametric regression setting, we extend \texttt{Rep} with a Lepski-style method that adapts to unknown structural features of the regression function.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-yu26b, title = { Undersmoothing Black-Box Models for Functional Estimation }, author = {Yu, Yue and Mukherjee, Debarghya and Banerjee, Moulinath and Ritov, Yaacov}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {2296--2304}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/yu26b/yu26b.pdf}, url = {https://proceedings.mlr.press/v300/yu26b.html}, abstract = { We study functional estimation using black-box models through a model-agnostic undersmoothing framework. The proposed procedure \texttt{Rep} operates by augmenting the original dataset through replicating a proportion of samples multiple times, and subsequently applying the black-box algorithm to the augmented dataset. This construction automatically induces undersmoothing and reduces the functional estimation error. We provide several empirical demonstrations (including neural network based learners) showing that compared to the plug-in estimator, the proposed algorithm \texttt{Rep} improves the estimation accuracy of functional estimation without requiring explicit expressions for the associated influence functions. Furthermore, we develop a theoretical analysis in two representative settings, the Nadaraya–Watson estimator and the random feature model, establishing that replication provides explicit prescriptions for the replication proportion and number of copies, and yields optimal convergence rates for functional estimation. In the classical nonparametric regression setting, we extend \texttt{Rep} with a Lepski-style method that adapts to unknown structural features of the regression function. } }
Endnote
%0 Conference Paper %T Undersmoothing Black-Box Models for Functional Estimation %A Yue Yu %A Debarghya Mukherjee %A Moulinath Banerjee %A Yaacov Ritov %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-yu26b %I PMLR %P 2296--2304 %U https://proceedings.mlr.press/v300/yu26b.html %V 300 %X We study functional estimation using black-box models through a model-agnostic undersmoothing framework. The proposed procedure \texttt{Rep} operates by augmenting the original dataset through replicating a proportion of samples multiple times, and subsequently applying the black-box algorithm to the augmented dataset. This construction automatically induces undersmoothing and reduces the functional estimation error. We provide several empirical demonstrations (including neural network based learners) showing that compared to the plug-in estimator, the proposed algorithm \texttt{Rep} improves the estimation accuracy of functional estimation without requiring explicit expressions for the associated influence functions. Furthermore, we develop a theoretical analysis in two representative settings, the Nadaraya–Watson estimator and the random feature model, establishing that replication provides explicit prescriptions for the replication proportion and number of copies, and yields optimal convergence rates for functional estimation. In the classical nonparametric regression setting, we extend \texttt{Rep} with a Lepski-style method that adapts to unknown structural features of the regression function.
APA
Yu, Y., Mukherjee, D., Banerjee, M. & Ritov, Y.. (2026). Undersmoothing Black-Box Models for Functional Estimation . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:2296-2304 Available from https://proceedings.mlr.press/v300/yu26b.html.

Related Material