Achieving Alignment Through Adaptive Play: Helping Optimize Objectives Without Observing Them

Jason T. Isa, Samuel Burden, Lillian J. Ratliff
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2336-2377, 2026.

Abstract

We study problems where two agents seek to minimize an objective that is known only to one of the agents. This setting arises in human-machine and {AI}-{AI} interactions where one agent’s objective is private knowledge inaccessible to the other, referred to as the "helper". We propose a game-theoretic learning algorithm that provably converges to optimal policies through repeated interaction without solving an inverse problem to recover the unknown objective. Importantly, the helper agent has no access to the cost function, its values, or its gradients, and operates under action-only feedback. Despite the information limitation, we establish convergence guarantees under {Polyak-Lojasiewicz} and Lipschitz-gradient assumptions. We validate the approach through two sets of experiments: human-machine interaction and a cart-pole environment with a reinforcement learning agent. In each of these experiments, the helper uses our proposed algorithm. Across scalar and multidimensional action spaces, we demonstrate consistent convergence under action-only feedback.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-isa26a, title = {Achieving Alignment Through Adaptive Play: Helping Optimize Objectives Without Observing Them}, author = {Isa, Jason T. and Burden, Samuel and Ratliff, Lillian J.}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {2336--2377}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/isa26a/isa26a.pdf}, url = {https://proceedings.mlr.press/v337/isa26a.html}, abstract = {We study problems where two agents seek to minimize an objective that is known only to one of the agents. This setting arises in human-machine and {AI}-{AI} interactions where one agent’s objective is private knowledge inaccessible to the other, referred to as the "helper". We propose a game-theoretic learning algorithm that provably converges to optimal policies through repeated interaction without solving an inverse problem to recover the unknown objective. Importantly, the helper agent has no access to the cost function, its values, or its gradients, and operates under action-only feedback. Despite the information limitation, we establish convergence guarantees under {Polyak-Lojasiewicz} and Lipschitz-gradient assumptions. We validate the approach through two sets of experiments: human-machine interaction and a cart-pole environment with a reinforcement learning agent. In each of these experiments, the helper uses our proposed algorithm. Across scalar and multidimensional action spaces, we demonstrate consistent convergence under action-only feedback.} }
Endnote
%0 Conference Paper %T Achieving Alignment Through Adaptive Play: Helping Optimize Objectives Without Observing Them %A Jason T. Isa %A Samuel Burden %A Lillian J. Ratliff %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-isa26a %I PMLR %P 2336--2377 %U https://proceedings.mlr.press/v337/isa26a.html %V 337 %X We study problems where two agents seek to minimize an objective that is known only to one of the agents. This setting arises in human-machine and {AI}-{AI} interactions where one agent’s objective is private knowledge inaccessible to the other, referred to as the "helper". We propose a game-theoretic learning algorithm that provably converges to optimal policies through repeated interaction without solving an inverse problem to recover the unknown objective. Importantly, the helper agent has no access to the cost function, its values, or its gradients, and operates under action-only feedback. Despite the information limitation, we establish convergence guarantees under {Polyak-Lojasiewicz} and Lipschitz-gradient assumptions. We validate the approach through two sets of experiments: human-machine interaction and a cart-pole environment with a reinforcement learning agent. In each of these experiments, the helper uses our proposed algorithm. Across scalar and multidimensional action spaces, we demonstrate consistent convergence under action-only feedback.
APA
Isa, J.T., Burden, S. & Ratliff, L.J.. (2026). Achieving Alignment Through Adaptive Play: Helping Optimize Objectives Without Observing Them. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:2336-2377 Available from https://proceedings.mlr.press/v337/isa26a.html.

Related Material