Importance Sampling for Fair Policy Selection

Shayan Doroudi, Philip Thomas, Emma Brunskill
Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, PMLR R15:431-440, 2017.

Abstract

We consider the problem of off-policy policy selection in reinforcement learning: using historical data generated from running one policy to compare two or more policies. We show that approaches based on importance sampling can be unfair—they can select the worse of the two policies more often than not. We give two examples where the unfairness of importance sampling could be practically concerning. We then present sufficient conditions to theoretically guarantee fairness and a related notion of safety. Finally, we provide a practical importance sampling-based estimator to help mitigate one of the systematic sources of unfairness resulting from using importance sampling for policy selection.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR15-doroudi17a, title = {Importance Sampling for Fair Policy Selection}, author = {Doroudi, Shayan and Thomas, Philip and Brunskill, Emma}, booktitle = {Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence}, pages = {431--440}, year = {2017}, editor = {Elidan, Gal and Kersting, Kristian}, volume = {R15}, series = {Proceedings of Machine Learning Research}, month = {11--15 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r15/main/assets/doroudi17a/doroudi17a.pdf}, url = {https://proceedings.mlr.press/r15/doroudi17a.html}, abstract = {We consider the problem of off-policy policy selection in reinforcement learning: using historical data generated from running one policy to compare two or more policies. We show that approaches based on importance sampling can be unfair—they can select the worse of the two policies more often than not. We give two examples where the unfairness of importance sampling could be practically concerning. We then present sufficient conditions to theoretically guarantee fairness and a related notion of safety. Finally, we provide a practical importance sampling-based estimator to help mitigate one of the systematic sources of unfairness resulting from using importance sampling for policy selection.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Importance Sampling for Fair Policy Selection %A Shayan Doroudi %A Philip Thomas %A Emma Brunskill %B Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2017 %E Gal Elidan %E Kristian Kersting %F pmlr-vR15-doroudi17a %I PMLR %P 431--440 %U https://proceedings.mlr.press/r15/doroudi17a.html %V R15 %X We consider the problem of off-policy policy selection in reinforcement learning: using historical data generated from running one policy to compare two or more policies. We show that approaches based on importance sampling can be unfair—they can select the worse of the two policies more often than not. We give two examples where the unfairness of importance sampling could be practically concerning. We then present sufficient conditions to theoretically guarantee fairness and a related notion of safety. Finally, we provide a practical importance sampling-based estimator to help mitigate one of the systematic sources of unfairness resulting from using importance sampling for policy selection. %Z Reissued by PMLR on 04 October 2026.
APA
Doroudi, S., Thomas, P. & Brunskill, E.. (2017). Importance Sampling for Fair Policy Selection. Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R15:431-440 Available from https://proceedings.mlr.press/r15/doroudi17a.html. Reissued by PMLR on 04 October 2026.

Related Material