Reflex: Real-Time Vision-Language-Action Control through Streaming Inference

Yuanchun Guo, Bingyan Liu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:38093-38108, 2026.

Abstract

Flow matching Vision-Language-Action (VLA) models promise precise continuous control, but their iterative denoising nature introduces fundamental incompatibilities with real-time robotics: global timestep injection invalidates KV-caching, forcing a choice between slow $O(N^2)$ re-computation or mathematically incorrect cache reuse. We present Reflex, a framework that enables real-time streaming inference for flow matching policies by exploiting the Timestep-Invariance Property—that perception encoders are functionally independent of the denoising loop. Reflex partitions the attention context into static, sliding, and dynamic regions, enabling $O(1)$ incremental cache updates while preserving full-batch-equivalent attention outputs for fixed inputs. To ensure stability under continuous high-frequency inference, we introduce AdaRMSNorm, an adaptive normalization layer that prevents BFloat16 numerical collapse by gating on flow phase. We further maximize throughput through an async pipeline that decouples visual encoding from action generation, combined with operator fusion that reduces kernel overhead. On LIBERO and Kinetix benchmarks, Reflex achieves a 2.58$\times$ inference speedup and 50Hz stable streaming, reducing reaction latency by up to 54% and enabling efficient deployment without performance degradation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-guo26l, title = {Reflex: Real-Time Vision-Language-Action Control through Streaming Inference}, author = {Guo, Yuanchun and Liu, Bingyan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {38093--38108}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/guo26l/guo26l.pdf}, url = {https://proceedings.mlr.press/v306/guo26l.html}, abstract = {Flow matching Vision-Language-Action (VLA) models promise precise continuous control, but their iterative denoising nature introduces fundamental incompatibilities with real-time robotics: global timestep injection invalidates KV-caching, forcing a choice between slow $O(N^2)$ re-computation or mathematically incorrect cache reuse. We present Reflex, a framework that enables real-time streaming inference for flow matching policies by exploiting the Timestep-Invariance Property—that perception encoders are functionally independent of the denoising loop. Reflex partitions the attention context into static, sliding, and dynamic regions, enabling $O(1)$ incremental cache updates while preserving full-batch-equivalent attention outputs for fixed inputs. To ensure stability under continuous high-frequency inference, we introduce AdaRMSNorm, an adaptive normalization layer that prevents BFloat16 numerical collapse by gating on flow phase. We further maximize throughput through an async pipeline that decouples visual encoding from action generation, combined with operator fusion that reduces kernel overhead. On LIBERO and Kinetix benchmarks, Reflex achieves a 2.58$\times$ inference speedup and 50Hz stable streaming, reducing reaction latency by up to 54% and enabling efficient deployment without performance degradation.} }
Endnote
%0 Conference Paper %T Reflex: Real-Time Vision-Language-Action Control through Streaming Inference %A Yuanchun Guo %A Bingyan Liu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-guo26l %I PMLR %P 38093--38108 %U https://proceedings.mlr.press/v306/guo26l.html %V 306 %X Flow matching Vision-Language-Action (VLA) models promise precise continuous control, but their iterative denoising nature introduces fundamental incompatibilities with real-time robotics: global timestep injection invalidates KV-caching, forcing a choice between slow $O(N^2)$ re-computation or mathematically incorrect cache reuse. We present Reflex, a framework that enables real-time streaming inference for flow matching policies by exploiting the Timestep-Invariance Property—that perception encoders are functionally independent of the denoising loop. Reflex partitions the attention context into static, sliding, and dynamic regions, enabling $O(1)$ incremental cache updates while preserving full-batch-equivalent attention outputs for fixed inputs. To ensure stability under continuous high-frequency inference, we introduce AdaRMSNorm, an adaptive normalization layer that prevents BFloat16 numerical collapse by gating on flow phase. We further maximize throughput through an async pipeline that decouples visual encoding from action generation, combined with operator fusion that reduces kernel overhead. On LIBERO and Kinetix benchmarks, Reflex achieves a 2.58$\times$ inference speedup and 50Hz stable streaming, reducing reaction latency by up to 54% and enabling efficient deployment without performance degradation.
APA
Guo, Y. & Liu, B.. (2026). Reflex: Real-Time Vision-Language-Action Control through Streaming Inference. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:38093-38108 Available from https://proceedings.mlr.press/v306/guo26l.html.

Related Material