[edit]
Learning-based Optimal Admission Control for Erlang-B Queuing Systems
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:6410-6444, 2026.
Abstract
This paper studies a learning-based optimal admission control policy for an Erlang–B queuing system under partial observation. Only at each arrival, the dispatcher observes the current occupancy and decides whether to accept or reject the job. A completed job yields a fixed reward but incurs a cost proportional to the service duration. The objective is to design an admission control policy that maximizes the long-term average reward. The arrival and service rates are unknown, and unobserved departures complicate parameter estimation. The reward structure induces an asymmetry in the optimal decision across parameter regimes. We establish an instance-dependent asymptotic lower bound, explicitly dependent on the system parameters, demonstrating that the worst-case regret is logarithmic in the number of arrivals in one regime.