Learning-based Optimal Admission Control for Erlang-B Queuing Systems

Shubhhi Singh, Shubhanshu Shekhar, Vijay G Subramanian
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:6410-6444, 2026.

Abstract

This paper studies a learning-based optimal admission control policy for an Erlang–B queuing system under partial observation. Only at each arrival, the dispatcher observes the current occupancy and decides whether to accept or reject the job. A completed job yields a fixed reward but incurs a cost proportional to the service duration. The objective is to design an admission control policy that maximizes the long-term average reward. The arrival and service rates are unknown, and unobserved departures complicate parameter estimation. The reward structure induces an asymmetry in the optimal decision across parameter regimes. We establish an instance-dependent asymptotic lower bound, explicitly dependent on the system parameters, demonstrating that the worst-case regret is logarithmic in the number of arrivals in one regime.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-singh26a, title = {Learning-based Optimal Admission Control for Erlang-B Queuing Systems}, author = {Singh, Shubhhi and Shekhar, Shubhanshu and Subramanian, Vijay G}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {6410--6444}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/singh26a/singh26a.pdf}, url = {https://proceedings.mlr.press/v337/singh26a.html}, abstract = {This paper studies a learning-based optimal admission control policy for an Erlang–B queuing system under partial observation. Only at each arrival, the dispatcher observes the current occupancy and decides whether to accept or reject the job. A completed job yields a fixed reward but incurs a cost proportional to the service duration. The objective is to design an admission control policy that maximizes the long-term average reward. The arrival and service rates are unknown, and unobserved departures complicate parameter estimation. The reward structure induces an asymmetry in the optimal decision across parameter regimes. We establish an instance-dependent asymptotic lower bound, explicitly dependent on the system parameters, demonstrating that the worst-case regret is logarithmic in the number of arrivals in one regime.} }
Endnote
%0 Conference Paper %T Learning-based Optimal Admission Control for Erlang-B Queuing Systems %A Shubhhi Singh %A Shubhanshu Shekhar %A Vijay G Subramanian %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-singh26a %I PMLR %P 6410--6444 %U https://proceedings.mlr.press/v337/singh26a.html %V 337 %X This paper studies a learning-based optimal admission control policy for an Erlang–B queuing system under partial observation. Only at each arrival, the dispatcher observes the current occupancy and decides whether to accept or reject the job. A completed job yields a fixed reward but incurs a cost proportional to the service duration. The objective is to design an admission control policy that maximizes the long-term average reward. The arrival and service rates are unknown, and unobserved departures complicate parameter estimation. The reward structure induces an asymmetry in the optimal decision across parameter regimes. We establish an instance-dependent asymptotic lower bound, explicitly dependent on the system parameters, demonstrating that the worst-case regret is logarithmic in the number of arrivals in one regime.
APA
Singh, S., Shekhar, S. & Subramanian, V.G.. (2026). Learning-based Optimal Admission Control for Erlang-B Queuing Systems. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:6410-6444 Available from https://proceedings.mlr.press/v337/singh26a.html.

Related Material