Learning Interpretable Options by Identifying Reward Diffusion Bottlenecks in Reinforcement Learning

Yiming Fei, Lang Qin, Rui Yan, Huajin Tang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:29804-29832, 2026.

Abstract

Bottleneck states, which connect distinct regions of the state space, provide a principled and interpretable basis for constructing temporal abstractions in Hierarchical Reinforcement Learning (HRL). However, existing bottleneck identification methods primarily rely on topological analysis of the state-transition graph, limiting their scalability to high-dimensional or continuous domains. To address this challenge, we introduce Value Power Strength (VPS), a value function-based metric inspired by the analogy between the Bellman equation and Kirchhoff’s current law, to quantify bottleneck property via the diffusion of reward in Markov Decision Processes (MDPs). VPS is estimated efficiently using value functions learned from random reward signals and captures reward diffusion bottlenecks in both discrete and continuous state spaces. Leveraging VPS, we design options that guide agents toward or away from bottleneck regions. Experiments on tabular domains, continuous-control, and Atari 2600 games show that VPS identifies semantically meaningful bottlenecks, while the learned options improve exploration.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-fei26a, title = {Learning Interpretable Options by Identifying Reward Diffusion Bottlenecks in Reinforcement Learning}, author = {Fei, Yiming and Qin, Lang and Yan, Rui and Tang, Huajin}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {29804--29832}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/fei26a/fei26a.pdf}, url = {https://proceedings.mlr.press/v306/fei26a.html}, abstract = {Bottleneck states, which connect distinct regions of the state space, provide a principled and interpretable basis for constructing temporal abstractions in Hierarchical Reinforcement Learning (HRL). However, existing bottleneck identification methods primarily rely on topological analysis of the state-transition graph, limiting their scalability to high-dimensional or continuous domains. To address this challenge, we introduce Value Power Strength (VPS), a value function-based metric inspired by the analogy between the Bellman equation and Kirchhoff’s current law, to quantify bottleneck property via the diffusion of reward in Markov Decision Processes (MDPs). VPS is estimated efficiently using value functions learned from random reward signals and captures reward diffusion bottlenecks in both discrete and continuous state spaces. Leveraging VPS, we design options that guide agents toward or away from bottleneck regions. Experiments on tabular domains, continuous-control, and Atari 2600 games show that VPS identifies semantically meaningful bottlenecks, while the learned options improve exploration.} }
Endnote
%0 Conference Paper %T Learning Interpretable Options by Identifying Reward Diffusion Bottlenecks in Reinforcement Learning %A Yiming Fei %A Lang Qin %A Rui Yan %A Huajin Tang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-fei26a %I PMLR %P 29804--29832 %U https://proceedings.mlr.press/v306/fei26a.html %V 306 %X Bottleneck states, which connect distinct regions of the state space, provide a principled and interpretable basis for constructing temporal abstractions in Hierarchical Reinforcement Learning (HRL). However, existing bottleneck identification methods primarily rely on topological analysis of the state-transition graph, limiting their scalability to high-dimensional or continuous domains. To address this challenge, we introduce Value Power Strength (VPS), a value function-based metric inspired by the analogy between the Bellman equation and Kirchhoff’s current law, to quantify bottleneck property via the diffusion of reward in Markov Decision Processes (MDPs). VPS is estimated efficiently using value functions learned from random reward signals and captures reward diffusion bottlenecks in both discrete and continuous state spaces. Leveraging VPS, we design options that guide agents toward or away from bottleneck regions. Experiments on tabular domains, continuous-control, and Atari 2600 games show that VPS identifies semantically meaningful bottlenecks, while the learned options improve exploration.
APA
Fei, Y., Qin, L., Yan, R. & Tang, H.. (2026). Learning Interpretable Options by Identifying Reward Diffusion Bottlenecks in Reinforcement Learning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:29804-29832 Available from https://proceedings.mlr.press/v306/fei26a.html.

Related Material