[edit]
Achieving Alignment Through Adaptive Play: Helping Optimize Objectives Without Observing Them
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2336-2377, 2026.
Abstract
We study problems where two agents seek to minimize an objective that is known only to one of the agents. This setting arises in human-machine and {AI}-{AI} interactions where one agent’s objective is private knowledge inaccessible to the other, referred to as the "helper". We propose a game-theoretic learning algorithm that provably converges to optimal policies through repeated interaction without solving an inverse problem to recover the unknown objective. Importantly, the helper agent has no access to the cost function, its values, or its gradients, and operates under action-only feedback. Despite the information limitation, we establish convergence guarantees under {Polyak-Lojasiewicz} and Lipschitz-gradient assumptions. We validate the approach through two sets of experiments: human-machine interaction and a cart-pole environment with a reinforcement learning agent. In each of these experiments, the helper uses our proposed algorithm. Across scalar and multidimensional action spaces, we demonstrate consistent convergence under action-only feedback.