[edit]
Neural Value Iteration
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7896-7912, 2026.
Abstract
The value function of a POMDP exhibits the piecewise-linear-convex (PWLC) property and can be represented as a finite set of hyperplanes, known as $\alpha$-vectors. Most state-of-the-art POMDP solvers (offline planners) follow the point-based value iteration scheme, which performs {Bellman} backups on $\alpha$-vectors at reachable belief points until convergence. However, since each $\alpha$-vector is $|S|$-dimensional, these methods quickly become intractable for large-scale problems due to the prohibitive computational cost of {Bellman} backups. In this work, we demonstrate that the PWLC representation of POMDP value functions can be generalized by replacing the alpha-vector sets with a finite set of neural networks. This insight enables a novel POMDP planning algorithm, called *Neural Value Iteration*, which combines the generalization capability of neural networks with the classical value iteration framework. Requiring only a black-box simulator, our approach scales to extremely large POMDPs that are intractable for existing offline solvers.