Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage

Siddhant Arora, Haidar Khan, Kai Sun, Xin Luna Dong, Sajal Choudhary, Seungwhan Moon, Xinyuan Zhang, Adithya Sagar, Surya Teja Appini, Kaushik Patnaik, Sanat Sharma, Shinji Watanabe, Anuj Kumar, Ahmed A Aly, Yue Liu, Florian Metze, Zhaojiang Lin
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:3633-3663, 2026.

Abstract

End-to-end speech-in, speech-out dialogue systems are emerging as a powerful alternative to traditional ASR–LLM–TTS pipelines but remain prone to hallucinations due to limited factual grounding. While text-based dialogue models have effectively mitigated this issue through tools such as web search APIs, extending such capabilities to speech-in, speech-out systems remains underexplored. A key challenge is that tool integration increases latency, disrupting conversational flow. To mitigate this, we propose Streaming Retrieval-Augmented Generation (Stream RAG), a novel framework that reduces latency by predicting tool queries in parallel with user speech, even before the user finishes speaking. Specifically, we develop a post-training pipeline that teaches the model when to issue tool calls and how to generate spoken summaries using retrieved text results, thereby improving both accuracy and responsiveness. To evaluate our approach, we construct AudioCRAG, a benchmark created by converting queries from the publicly available CRAG dataset into speech form. Experimental results show that Stream RAG improves QA accuracy by over 20.0% absolute on AudioCRAG and achieves state-of-the-art performance, including outperforming cascaded systems, on the SLUE-SQA benchmark, while reducing latency by up to 57%. Stream RAG is modality-agnostic and can be applied equally to typed input, paving the way for more agentic, real-time AI assistants.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-arora26a, title = {Stream {RAG}: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage}, author = {Arora, Siddhant and Khan, Haidar and Sun, Kai and Dong, Xin Luna and Choudhary, Sajal and Moon, Seungwhan and Zhang, Xinyuan and Sagar, Adithya and Appini, Surya Teja and Patnaik, Kaushik and Sharma, Sanat and Watanabe, Shinji and Kumar, Anuj and Aly, Ahmed A and Liu, Yue and Metze, Florian and Lin, Zhaojiang}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {3633--3663}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/arora26a/arora26a.pdf}, url = {https://proceedings.mlr.press/v306/arora26a.html}, abstract = {End-to-end speech-in, speech-out dialogue systems are emerging as a powerful alternative to traditional ASR–LLM–TTS pipelines but remain prone to hallucinations due to limited factual grounding. While text-based dialogue models have effectively mitigated this issue through tools such as web search APIs, extending such capabilities to speech-in, speech-out systems remains underexplored. A key challenge is that tool integration increases latency, disrupting conversational flow. To mitigate this, we propose Streaming Retrieval-Augmented Generation (Stream RAG), a novel framework that reduces latency by predicting tool queries in parallel with user speech, even before the user finishes speaking. Specifically, we develop a post-training pipeline that teaches the model when to issue tool calls and how to generate spoken summaries using retrieved text results, thereby improving both accuracy and responsiveness. To evaluate our approach, we construct AudioCRAG, a benchmark created by converting queries from the publicly available CRAG dataset into speech form. Experimental results show that Stream RAG improves QA accuracy by over 20.0% absolute on AudioCRAG and achieves state-of-the-art performance, including outperforming cascaded systems, on the SLUE-SQA benchmark, while reducing latency by up to 57%. Stream RAG is modality-agnostic and can be applied equally to typed input, paving the way for more agentic, real-time AI assistants.} }
Endnote
%0 Conference Paper %T Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage %A Siddhant Arora %A Haidar Khan %A Kai Sun %A Xin Luna Dong %A Sajal Choudhary %A Seungwhan Moon %A Xinyuan Zhang %A Adithya Sagar %A Surya Teja Appini %A Kaushik Patnaik %A Sanat Sharma %A Shinji Watanabe %A Anuj Kumar %A Ahmed A Aly %A Yue Liu %A Florian Metze %A Zhaojiang Lin %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-arora26a %I PMLR %P 3633--3663 %U https://proceedings.mlr.press/v306/arora26a.html %V 306 %X End-to-end speech-in, speech-out dialogue systems are emerging as a powerful alternative to traditional ASR–LLM–TTS pipelines but remain prone to hallucinations due to limited factual grounding. While text-based dialogue models have effectively mitigated this issue through tools such as web search APIs, extending such capabilities to speech-in, speech-out systems remains underexplored. A key challenge is that tool integration increases latency, disrupting conversational flow. To mitigate this, we propose Streaming Retrieval-Augmented Generation (Stream RAG), a novel framework that reduces latency by predicting tool queries in parallel with user speech, even before the user finishes speaking. Specifically, we develop a post-training pipeline that teaches the model when to issue tool calls and how to generate spoken summaries using retrieved text results, thereby improving both accuracy and responsiveness. To evaluate our approach, we construct AudioCRAG, a benchmark created by converting queries from the publicly available CRAG dataset into speech form. Experimental results show that Stream RAG improves QA accuracy by over 20.0% absolute on AudioCRAG and achieves state-of-the-art performance, including outperforming cascaded systems, on the SLUE-SQA benchmark, while reducing latency by up to 57%. Stream RAG is modality-agnostic and can be applied equally to typed input, paving the way for more agentic, real-time AI assistants.
APA
Arora, S., Khan, H., Sun, K., Dong, X.L., Choudhary, S., Moon, S., Zhang, X., Sagar, A., Appini, S.T., Patnaik, K., Sharma, S., Watanabe, S., Kumar, A., Aly, A.A., Liu, Y., Metze, F. & Lin, Z.. (2026). Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:3633-3663 Available from https://proceedings.mlr.press/v306/arora26a.html.

Related Material