[edit]
$f$-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:29700-29733, 2026.
Abstract
In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities centered at the batch mean, i.e. variance of the difference in logprobs, is an effective, low variance, surrogate loss for training generative models. This loss has the property that when evaluated on-policy its gradients correspond to those of the KL divergence, while off-policy it remains a valid loss with the same global minimizer. Mean centreing the difference guarentees that the loss is valid for unnormalized target distributions, with variants of this being applied to large-scale RL tuning LLMs in KIMI K2. In this work, building on recent theoretical equivalences established in the GFlowNet literature, we show that this construction extends to the whole family of $f$-divergences. Specifically, utilizing an established one-to-one correspondence between translation invariant loss functions and $f$-divergences, we adapt this framework to policy optimization with unnormalized targets in a batch-wise fashion. This equivalence allows us to design new surrogate loss functions for tuning a wide class of generative models that are low variance and inherit the properties of the corresponding $f$-divergence, such as being more mode covering, whilst being applicable to off-policy data. We apply our losses on a range of tasks, including SynFlowNets for molecule discovery, conditional sampling of diffusion models, and asynchronous large language model (LLM) tuning, demonstrating that our models retain their predicted properties on- and off-policy in a wide class of generative models.