Kernel Treatment Effects with Adaptively Collected Data

Houssam Zenati, Bariscan Bozkurt, Arthur Gretton
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:3412-3420, 2026.

Abstract

Adaptive experiments improve efficiency by adjusting treatment assignments based on past outcomes, but this adaptivity breaks the i.i.d. assumptions that underpin classical asymptotics. At the same time, many questions of interest are distributional, extending beyond average effects. Kernel treatment effects (KTE) provide a flexible framework by representing interventional outcome distributions in an RKHS and comparing them via kernel distances. We present the first kernel-based framework for distributional inference under adaptive data collection. Our method combines doubly robust RKHS scores with a witness function learned on one fold, and performs inference on a second fold using a projected, sequentially normalized scalar statistic with valid type-I error. Experiments show that the resulting procedure is well calibrated and effective for both mean shifts and higher-moment differences, outperforming adaptive baselines limited to scalar effects.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-zenati26a, title = { Kernel Treatment Effects with Adaptively Collected Data }, author = {Zenati, Houssam and Bozkurt, Bariscan and Gretton, Arthur}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {3412--3420}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/zenati26a/zenati26a.pdf}, url = {https://proceedings.mlr.press/v300/zenati26a.html}, abstract = { Adaptive experiments improve efficiency by adjusting treatment assignments based on past outcomes, but this adaptivity breaks the i.i.d. assumptions that underpin classical asymptotics. At the same time, many questions of interest are distributional, extending beyond average effects. Kernel treatment effects (KTE) provide a flexible framework by representing interventional outcome distributions in an RKHS and comparing them via kernel distances. We present the first kernel-based framework for distributional inference under adaptive data collection. Our method combines doubly robust RKHS scores with a witness function learned on one fold, and performs inference on a second fold using a projected, sequentially normalized scalar statistic with valid type-I error. Experiments show that the resulting procedure is well calibrated and effective for both mean shifts and higher-moment differences, outperforming adaptive baselines limited to scalar effects. } }
Endnote
%0 Conference Paper %T Kernel Treatment Effects with Adaptively Collected Data %A Houssam Zenati %A Bariscan Bozkurt %A Arthur Gretton %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-zenati26a %I PMLR %P 3412--3420 %U https://proceedings.mlr.press/v300/zenati26a.html %V 300 %X Adaptive experiments improve efficiency by adjusting treatment assignments based on past outcomes, but this adaptivity breaks the i.i.d. assumptions that underpin classical asymptotics. At the same time, many questions of interest are distributional, extending beyond average effects. Kernel treatment effects (KTE) provide a flexible framework by representing interventional outcome distributions in an RKHS and comparing them via kernel distances. We present the first kernel-based framework for distributional inference under adaptive data collection. Our method combines doubly robust RKHS scores with a witness function learned on one fold, and performs inference on a second fold using a projected, sequentially normalized scalar statistic with valid type-I error. Experiments show that the resulting procedure is well calibrated and effective for both mean shifts and higher-moment differences, outperforming adaptive baselines limited to scalar effects.
APA
Zenati, H., Bozkurt, B. & Gretton, A.. (2026). Kernel Treatment Effects with Adaptively Collected Data . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:3412-3420 Available from https://proceedings.mlr.press/v300/zenati26a.html.

Related Material