A Model-Free Universal AI

Yegon Kim, Juho Lee
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3012-3035, 2026.

Abstract

In general reinforcement learning, all established optimal agents, including {AIXI}, are model-based, explicitly maintaining and using environment models. This paper introduces Universal {AI} with {Q-Induction} ({AIQI}), the first model-free agent proven to be asymptotically $\varepsilon$-optimal in general {RL}. {AIQI} performs universal induction over distributional action-value functions, instead of policies or environments like previous works. Under a grain of truth condition, we prove that {AIQI} is strong asymptotically $\varepsilon$-optimal and asymptotically $\varepsilon$-{Bayes}-optimal. We also apply our novel proof techniques to show asymptotic $\varepsilon$-optimality of Self-{AIXI} without any ad-hoc assumptions. Our results significantly expand the diversity of known universal agents.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-kim26c, title = {A Model-Free Universal {AI}}, author = {Kim, Yegon and Lee, Juho}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {3012--3035}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/kim26c/kim26c.pdf}, url = {https://proceedings.mlr.press/v337/kim26c.html}, abstract = {In general reinforcement learning, all established optimal agents, including {AIXI}, are model-based, explicitly maintaining and using environment models. This paper introduces Universal {AI} with {Q-Induction} ({AIQI}), the first model-free agent proven to be asymptotically $\varepsilon$-optimal in general {RL}. {AIQI} performs universal induction over distributional action-value functions, instead of policies or environments like previous works. Under a grain of truth condition, we prove that {AIQI} is strong asymptotically $\varepsilon$-optimal and asymptotically $\varepsilon$-{Bayes}-optimal. We also apply our novel proof techniques to show asymptotic $\varepsilon$-optimality of Self-{AIXI} without any ad-hoc assumptions. Our results significantly expand the diversity of known universal agents.} }
Endnote
%0 Conference Paper %T A Model-Free Universal AI %A Yegon Kim %A Juho Lee %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-kim26c %I PMLR %P 3012--3035 %U https://proceedings.mlr.press/v337/kim26c.html %V 337 %X In general reinforcement learning, all established optimal agents, including {AIXI}, are model-based, explicitly maintaining and using environment models. This paper introduces Universal {AI} with {Q-Induction} ({AIQI}), the first model-free agent proven to be asymptotically $\varepsilon$-optimal in general {RL}. {AIQI} performs universal induction over distributional action-value functions, instead of policies or environments like previous works. Under a grain of truth condition, we prove that {AIQI} is strong asymptotically $\varepsilon$-optimal and asymptotically $\varepsilon$-{Bayes}-optimal. We also apply our novel proof techniques to show asymptotic $\varepsilon$-optimality of Self-{AIXI} without any ad-hoc assumptions. Our results significantly expand the diversity of known universal agents.
APA
Kim, Y. & Lee, J.. (2026). A Model-Free Universal AI. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:3012-3035 Available from https://proceedings.mlr.press/v337/kim26c.html.

Related Material