Learning in Continuous State-Space MDPs for Network Inventory Management

Hansheng Jiang, Shunan Jiang, Zuo-Jun Shen
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:667-675, 2026.

Abstract

We consider online learning in infinite-horizon, average-cost Markov Decision Processes (MDPs) with multi-dimensional, continuous state spaces and censored feedback. Our model setting, motivated by network inventory management applications such as vehicle sharing, is characterized by complex, correlated state transitions and the absence of value function convexity, rendering standard analytical techniques for both MDPs and inventory control inapplicable. Our primary contribution is an integrated framework establishing and leveraging the Lipschitz property of the long-run average cost function. This insight allows us to analyze the problem through the lens of Lipschitz bandits, for which we design a provably efficient online learning algorithm that learns a near-optimal policy from censored demand data. We derive a high-probability regret bound of $O(T^{\frac{n}{n+1}} (\log T)^{\frac{1}{n+1}})$, where $n$ is the network size through customized concentration inequalities for cumulative costs in MDPs with state-dependent transitions. Furthermore, we devise a matching lower bound for this learning problem, which captures the inherent dimensionality challenge.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-jiang26a, title = { Learning in Continuous State-Space MDPs for Network Inventory Management }, author = {Jiang, Hansheng and Jiang, Shunan and Shen, Zuo-Jun}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {667--675}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/jiang26a/jiang26a.pdf}, url = {https://proceedings.mlr.press/v300/jiang26a.html}, abstract = { We consider online learning in infinite-horizon, average-cost Markov Decision Processes (MDPs) with multi-dimensional, continuous state spaces and censored feedback. Our model setting, motivated by network inventory management applications such as vehicle sharing, is characterized by complex, correlated state transitions and the absence of value function convexity, rendering standard analytical techniques for both MDPs and inventory control inapplicable. Our primary contribution is an integrated framework establishing and leveraging the Lipschitz property of the long-run average cost function. This insight allows us to analyze the problem through the lens of Lipschitz bandits, for which we design a provably efficient online learning algorithm that learns a near-optimal policy from censored demand data. We derive a high-probability regret bound of $O(T^{\frac{n}{n+1}} (\log T)^{\frac{1}{n+1}})$, where $n$ is the network size through customized concentration inequalities for cumulative costs in MDPs with state-dependent transitions. Furthermore, we devise a matching lower bound for this learning problem, which captures the inherent dimensionality challenge. } }
Endnote
%0 Conference Paper %T Learning in Continuous State-Space MDPs for Network Inventory Management %A Hansheng Jiang %A Shunan Jiang %A Zuo-Jun Shen %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-jiang26a %I PMLR %P 667--675 %U https://proceedings.mlr.press/v300/jiang26a.html %V 300 %X We consider online learning in infinite-horizon, average-cost Markov Decision Processes (MDPs) with multi-dimensional, continuous state spaces and censored feedback. Our model setting, motivated by network inventory management applications such as vehicle sharing, is characterized by complex, correlated state transitions and the absence of value function convexity, rendering standard analytical techniques for both MDPs and inventory control inapplicable. Our primary contribution is an integrated framework establishing and leveraging the Lipschitz property of the long-run average cost function. This insight allows us to analyze the problem through the lens of Lipschitz bandits, for which we design a provably efficient online learning algorithm that learns a near-optimal policy from censored demand data. We derive a high-probability regret bound of $O(T^{\frac{n}{n+1}} (\log T)^{\frac{1}{n+1}})$, where $n$ is the network size through customized concentration inequalities for cumulative costs in MDPs with state-dependent transitions. Furthermore, we devise a matching lower bound for this learning problem, which captures the inherent dimensionality challenge.
APA
Jiang, H., Jiang, S. & Shen, Z.. (2026). Learning in Continuous State-Space MDPs for Network Inventory Management . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:667-675 Available from https://proceedings.mlr.press/v300/jiang26a.html.

Related Material