Active Measuring in Reinforcement Learning With Delayed Negative Effects

Daiqi Gao, Ziping Xu, Aseel Rawashdeh, Predrag Klasnja, Susan Murphy
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:3907-3915, 2026.

Abstract

Measuring states in reinforcement learning (RL) can be costly in real-world settings and may negatively influence future outcomes. We introduce the Actively Observable Markov Decision Process (AOMDP), where an agent not only selects control actions but also decides whether to measure the latent state. The measurement action reveals the true latent state but may have a negative delayed effect on the environment. We show that this reduced uncertainty enables sample-efficient learning and may increase the value of the optimal policy despite these costs. We formulate an AOMDP as a periodic partially observable MDP and propose an online RL algorithm based on belief states. To approximate the belief states, we further propose a sequential Monte Carlo method to jointly approximate the posterior of unknown static environment parameters and unobserved latent states. We evaluate the proposed algorithm in a digital health application, where the agent decides when to deliver digital interventions and when to assess users’ psychological status through surveys.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-gao26d, title = { Active Measuring in Reinforcement Learning With Delayed Negative Effects }, author = {Gao, Daiqi and Xu, Ziping and Rawashdeh, Aseel and Klasnja, Predrag and Murphy, Susan}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {3907--3915}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/gao26d/gao26d.pdf}, url = {https://proceedings.mlr.press/v300/gao26d.html}, abstract = { Measuring states in reinforcement learning (RL) can be costly in real-world settings and may negatively influence future outcomes. We introduce the Actively Observable Markov Decision Process (AOMDP), where an agent not only selects control actions but also decides whether to measure the latent state. The measurement action reveals the true latent state but may have a negative delayed effect on the environment. We show that this reduced uncertainty enables sample-efficient learning and may increase the value of the optimal policy despite these costs. We formulate an AOMDP as a periodic partially observable MDP and propose an online RL algorithm based on belief states. To approximate the belief states, we further propose a sequential Monte Carlo method to jointly approximate the posterior of unknown static environment parameters and unobserved latent states. We evaluate the proposed algorithm in a digital health application, where the agent decides when to deliver digital interventions and when to assess users’ psychological status through surveys. } }
Endnote
%0 Conference Paper %T Active Measuring in Reinforcement Learning With Delayed Negative Effects %A Daiqi Gao %A Ziping Xu %A Aseel Rawashdeh %A Predrag Klasnja %A Susan Murphy %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-gao26d %I PMLR %P 3907--3915 %U https://proceedings.mlr.press/v300/gao26d.html %V 300 %X Measuring states in reinforcement learning (RL) can be costly in real-world settings and may negatively influence future outcomes. We introduce the Actively Observable Markov Decision Process (AOMDP), where an agent not only selects control actions but also decides whether to measure the latent state. The measurement action reveals the true latent state but may have a negative delayed effect on the environment. We show that this reduced uncertainty enables sample-efficient learning and may increase the value of the optimal policy despite these costs. We formulate an AOMDP as a periodic partially observable MDP and propose an online RL algorithm based on belief states. To approximate the belief states, we further propose a sequential Monte Carlo method to jointly approximate the posterior of unknown static environment parameters and unobserved latent states. We evaluate the proposed algorithm in a digital health application, where the agent decides when to deliver digital interventions and when to assess users’ psychological status through surveys.
APA
Gao, D., Xu, Z., Rawashdeh, A., Klasnja, P. & Murphy, S.. (2026). Active Measuring in Reinforcement Learning With Delayed Negative Effects . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:3907-3915 Available from https://proceedings.mlr.press/v300/gao26d.html.

Related Material