[edit]
POMDPs under Probabilistic Semantics
Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence, PMLR R11:65-74, 2013.
Abstract
We consider partially observable Markov decision processes (POMDPs) with limit- average payoff, where a reward value in the interval [0, 1] is associated to every transi- tion, and the payoffof an infinite path is the long-run average of the rewards. We con- sider two types of path constraints: (i) quan- titative constraint defines the set of paths where the payoffis at least a given thresh- old $\lambda$1 $\in$(0, 1]; and (ii) qualitative constraint which is a special case of quantitative con- straint with $\lambda$1 = 1. We consider the compu- tation of the almost-sure winning set, where the controller needs to ensure that the path constraint is satisfied with probability 1. Our main results for qualitative path constraint are as follows: (i) the problem of deciding the existence of a finite-memory controller is EXPTIME-complete; and (ii) the problem of deciding the existence of an infinite-memory controller is undecidable. For quantitative path constraint we show that the problem of deciding the existence of a finite-memory controller is undecidable.