$f$-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data

Jake Fawkes, Jason Hartford
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:29700-29733, 2026.

Abstract

In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities centered at the batch mean, i.e. variance of the difference in logprobs, is an effective, low variance, surrogate loss for training generative models. This loss has the property that when evaluated on-policy its gradients correspond to those of the KL divergence, while off-policy it remains a valid loss with the same global minimizer. Mean centreing the difference guarentees that the loss is valid for unnormalized target distributions, with variants of this being applied to large-scale RL tuning LLMs in KIMI K2. In this work, building on recent theoretical equivalences established in the GFlowNet literature, we show that this construction extends to the whole family of $f$-divergences. Specifically, utilizing an established one-to-one correspondence between translation invariant loss functions and $f$-divergences, we adapt this framework to policy optimization with unnormalized targets in a batch-wise fashion. This equivalence allows us to design new surrogate loss functions for tuning a wide class of generative models that are low variance and inherit the properties of the corresponding $f$-divergence, such as being more mode covering, whilst being applicable to off-policy data. We apply our losses on a range of tasks, including SynFlowNets for molecule discovery, conditional sampling of diffusion models, and asynchronous large language model (LLM) tuning, demonstrating that our models retain their predicted properties on- and off-policy in a wide class of generative models.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-fawkes26a, title = {$f$-Trajectory Balance: A Loss Family for Tuning {GF}low{N}ets, Generative Models, and {LLM}s with Off- and On-Policy Data}, author = {Fawkes, Jake and Hartford, Jason}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {29700--29733}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/fawkes26a/fawkes26a.pdf}, url = {https://proceedings.mlr.press/v306/fawkes26a.html}, abstract = {In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities centered at the batch mean, i.e. variance of the difference in logprobs, is an effective, low variance, surrogate loss for training generative models. This loss has the property that when evaluated on-policy its gradients correspond to those of the KL divergence, while off-policy it remains a valid loss with the same global minimizer. Mean centreing the difference guarentees that the loss is valid for unnormalized target distributions, with variants of this being applied to large-scale RL tuning LLMs in KIMI K2. In this work, building on recent theoretical equivalences established in the GFlowNet literature, we show that this construction extends to the whole family of $f$-divergences. Specifically, utilizing an established one-to-one correspondence between translation invariant loss functions and $f$-divergences, we adapt this framework to policy optimization with unnormalized targets in a batch-wise fashion. This equivalence allows us to design new surrogate loss functions for tuning a wide class of generative models that are low variance and inherit the properties of the corresponding $f$-divergence, such as being more mode covering, whilst being applicable to off-policy data. We apply our losses on a range of tasks, including SynFlowNets for molecule discovery, conditional sampling of diffusion models, and asynchronous large language model (LLM) tuning, demonstrating that our models retain their predicted properties on- and off-policy in a wide class of generative models.} }
Endnote
%0 Conference Paper %T $f$-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data %A Jake Fawkes %A Jason Hartford %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-fawkes26a %I PMLR %P 29700--29733 %U https://proceedings.mlr.press/v306/fawkes26a.html %V 306 %X In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities centered at the batch mean, i.e. variance of the difference in logprobs, is an effective, low variance, surrogate loss for training generative models. This loss has the property that when evaluated on-policy its gradients correspond to those of the KL divergence, while off-policy it remains a valid loss with the same global minimizer. Mean centreing the difference guarentees that the loss is valid for unnormalized target distributions, with variants of this being applied to large-scale RL tuning LLMs in KIMI K2. In this work, building on recent theoretical equivalences established in the GFlowNet literature, we show that this construction extends to the whole family of $f$-divergences. Specifically, utilizing an established one-to-one correspondence between translation invariant loss functions and $f$-divergences, we adapt this framework to policy optimization with unnormalized targets in a batch-wise fashion. This equivalence allows us to design new surrogate loss functions for tuning a wide class of generative models that are low variance and inherit the properties of the corresponding $f$-divergence, such as being more mode covering, whilst being applicable to off-policy data. We apply our losses on a range of tasks, including SynFlowNets for molecule discovery, conditional sampling of diffusion models, and asynchronous large language model (LLM) tuning, demonstrating that our models retain their predicted properties on- and off-policy in a wide class of generative models.
APA
Fawkes, J. & Hartford, J.. (2026). $f$-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:29700-29733 Available from https://proceedings.mlr.press/v306/fawkes26a.html.

Related Material