Test-Time Reinforcement Learning for Flow Matching

Jili Chen, Changqin Huang, Qionghao Huang, Yaxin Tu, Zhonglong Zheng, Xiaodi Huang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:14589-14624, 2026.

Abstract

Flow-matching has emerged as a leading framework for high-fidelity text-to-image generation. However, its alignment with human preferences through RL is often hindered by substantial computational overhead. In this paper, we introduce Flow-TTRL, the first test-time reinforcement learning framework that achieves alignment on the fly. Our approach reinterprets intermediate latent representations as an implicit policy and utilizes SDE-based rollouts to explore high-reward trajectories within the learned vector field. Specifically, we propose a two-stage optimization strategy: Proximal Reward Difference Prediction (PRDP) ensures structural stability in high-noise regimes through pairwise reward regression, while Group Relative Policy Optimization (GRPO) refines fine-grained aesthetic details by maximizing relative advantages within sampled candidate groups. Experimental results show that Flow-TTRL significantly boosts aesthetic quality, text-image alignment, and human preference across diverse backbones. On the GenEval benchmark, Flow-TTRL elevates the accuracy of SD 3.5-Medium from 63% to 87% and Flux.1 Dev from 66% to 83%. Furthermore, our framework achieves an average gain of 15% to 20% across T2I-CompBench metrics, delivering performance comparable to state-of-the-art RL-based fine-tuning methods without the need for additional fine-tuning. Our code is available at https://github.com/TheShy-Dream/Flow-TTRL.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26as, title = {Test-Time Reinforcement Learning for Flow Matching}, author = {Chen, Jili and Huang, Changqin and Huang, Qionghao and Tu, Yaxin and Zheng, Zhonglong and Huang, Xiaodi}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {14589--14624}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26as/chen26as.pdf}, url = {https://proceedings.mlr.press/v306/chen26as.html}, abstract = {Flow-matching has emerged as a leading framework for high-fidelity text-to-image generation. However, its alignment with human preferences through RL is often hindered by substantial computational overhead. In this paper, we introduce Flow-TTRL, the first test-time reinforcement learning framework that achieves alignment on the fly. Our approach reinterprets intermediate latent representations as an implicit policy and utilizes SDE-based rollouts to explore high-reward trajectories within the learned vector field. Specifically, we propose a two-stage optimization strategy: Proximal Reward Difference Prediction (PRDP) ensures structural stability in high-noise regimes through pairwise reward regression, while Group Relative Policy Optimization (GRPO) refines fine-grained aesthetic details by maximizing relative advantages within sampled candidate groups. Experimental results show that Flow-TTRL significantly boosts aesthetic quality, text-image alignment, and human preference across diverse backbones. On the GenEval benchmark, Flow-TTRL elevates the accuracy of SD 3.5-Medium from 63% to 87% and Flux.1 Dev from 66% to 83%. Furthermore, our framework achieves an average gain of 15% to 20% across T2I-CompBench metrics, delivering performance comparable to state-of-the-art RL-based fine-tuning methods without the need for additional fine-tuning. Our code is available at https://github.com/TheShy-Dream/Flow-TTRL.} }
Endnote
%0 Conference Paper %T Test-Time Reinforcement Learning for Flow Matching %A Jili Chen %A Changqin Huang %A Qionghao Huang %A Yaxin Tu %A Zhonglong Zheng %A Xiaodi Huang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26as %I PMLR %P 14589--14624 %U https://proceedings.mlr.press/v306/chen26as.html %V 306 %X Flow-matching has emerged as a leading framework for high-fidelity text-to-image generation. However, its alignment with human preferences through RL is often hindered by substantial computational overhead. In this paper, we introduce Flow-TTRL, the first test-time reinforcement learning framework that achieves alignment on the fly. Our approach reinterprets intermediate latent representations as an implicit policy and utilizes SDE-based rollouts to explore high-reward trajectories within the learned vector field. Specifically, we propose a two-stage optimization strategy: Proximal Reward Difference Prediction (PRDP) ensures structural stability in high-noise regimes through pairwise reward regression, while Group Relative Policy Optimization (GRPO) refines fine-grained aesthetic details by maximizing relative advantages within sampled candidate groups. Experimental results show that Flow-TTRL significantly boosts aesthetic quality, text-image alignment, and human preference across diverse backbones. On the GenEval benchmark, Flow-TTRL elevates the accuracy of SD 3.5-Medium from 63% to 87% and Flux.1 Dev from 66% to 83%. Furthermore, our framework achieves an average gain of 15% to 20% across T2I-CompBench metrics, delivering performance comparable to state-of-the-art RL-based fine-tuning methods without the need for additional fine-tuning. Our code is available at https://github.com/TheShy-Dream/Flow-TTRL.
APA
Chen, J., Huang, C., Huang, Q., Tu, Y., Zheng, Z. & Huang, X.. (2026). Test-Time Reinforcement Learning for Flow Matching. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:14589-14624 Available from https://proceedings.mlr.press/v306/chen26as.html.

Related Material