Learning Anisotropic Value Geometry with Finsler Reinforcement Learning

Jumman Hossain, Nirmalya Roy
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:43950-43983, 2026.

Abstract

We introduce Finslerian Reinforcement Learning (FiRL), an RL framework that makes directional costs explicit and improves robustness to tail risk. FiRL incorporates a Finsler metric into the locomotion cost, expressing effort as $F(x,v)$ that depends on the state $x$ and motion $v$, so it can capture uphill versus downhill asymmetry, lateral slip, and other direction-dependent effects. To handle rare but catastrophic outcomes, FiRL optimizes a Conditional Value-at-Risk ($CVaR_\alpha$) objective. We derive the corresponding risk-sensitive Bellman equation and show that the resulting CVaR–Finsler Bellman operator is a $\gamma$-contraction. This guarantees a unique fixed-point value function, while the underlying Finsler cost induces an asymmetric path cost $d_F$ that satisfies a triangle inequality despite directional asymmetry. We then develop a FiRL actor–critic algorithm to learn policies under this anisotropic, risk-averse objective. Across simulation benchmarks and real-world robot trials, FiRL demonstrates safer and more energy-efficient locomotion behavior than strong baselines such as risk-neutral PPO. For instance, on a $12^\circ$ sloped Hopper task, FiRL reduces worst-case ($CVaR_{0.1}$) impact forces by over 35% and total energy cost by 15%, while also improving success rate.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-hossain26a, title = {Learning Anisotropic Value Geometry with Finsler Reinforcement Learning}, author = {Hossain, Jumman and Roy, Nirmalya}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {43950--43983}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/hossain26a/hossain26a.pdf}, url = {https://proceedings.mlr.press/v306/hossain26a.html}, abstract = {We introduce Finslerian Reinforcement Learning (FiRL), an RL framework that makes directional costs explicit and improves robustness to tail risk. FiRL incorporates a Finsler metric into the locomotion cost, expressing effort as $F(x,v)$ that depends on the state $x$ and motion $v$, so it can capture uphill versus downhill asymmetry, lateral slip, and other direction-dependent effects. To handle rare but catastrophic outcomes, FiRL optimizes a Conditional Value-at-Risk ($CVaR_\alpha$) objective. We derive the corresponding risk-sensitive Bellman equation and show that the resulting CVaR–Finsler Bellman operator is a $\gamma$-contraction. This guarantees a unique fixed-point value function, while the underlying Finsler cost induces an asymmetric path cost $d_F$ that satisfies a triangle inequality despite directional asymmetry. We then develop a FiRL actor–critic algorithm to learn policies under this anisotropic, risk-averse objective. Across simulation benchmarks and real-world robot trials, FiRL demonstrates safer and more energy-efficient locomotion behavior than strong baselines such as risk-neutral PPO. For instance, on a $12^\circ$ sloped Hopper task, FiRL reduces worst-case ($CVaR_{0.1}$) impact forces by over 35% and total energy cost by 15%, while also improving success rate.} }
Endnote
%0 Conference Paper %T Learning Anisotropic Value Geometry with Finsler Reinforcement Learning %A Jumman Hossain %A Nirmalya Roy %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-hossain26a %I PMLR %P 43950--43983 %U https://proceedings.mlr.press/v306/hossain26a.html %V 306 %X We introduce Finslerian Reinforcement Learning (FiRL), an RL framework that makes directional costs explicit and improves robustness to tail risk. FiRL incorporates a Finsler metric into the locomotion cost, expressing effort as $F(x,v)$ that depends on the state $x$ and motion $v$, so it can capture uphill versus downhill asymmetry, lateral slip, and other direction-dependent effects. To handle rare but catastrophic outcomes, FiRL optimizes a Conditional Value-at-Risk ($CVaR_\alpha$) objective. We derive the corresponding risk-sensitive Bellman equation and show that the resulting CVaR–Finsler Bellman operator is a $\gamma$-contraction. This guarantees a unique fixed-point value function, while the underlying Finsler cost induces an asymmetric path cost $d_F$ that satisfies a triangle inequality despite directional asymmetry. We then develop a FiRL actor–critic algorithm to learn policies under this anisotropic, risk-averse objective. Across simulation benchmarks and real-world robot trials, FiRL demonstrates safer and more energy-efficient locomotion behavior than strong baselines such as risk-neutral PPO. For instance, on a $12^\circ$ sloped Hopper task, FiRL reduces worst-case ($CVaR_{0.1}$) impact forces by over 35% and total energy cost by 15%, while also improving success rate.
APA
Hossain, J. & Roy, N.. (2026). Learning Anisotropic Value Geometry with Finsler Reinforcement Learning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:43950-43983 Available from https://proceedings.mlr.press/v306/hossain26a.html.

Related Material