DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use

Aili Chen, Chi Zhang, Junteng Liu, Jiangjie Chen, Chengyu Du, Yunji Li, Ming Zhong, Qin Wang, Zhengmao Zhu, Jiayuan Song, Ke Ji, Junxian He, Pengyu Zhao, Yanghua Xiao
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:18235-18258, 2026.

Abstract

Recent work increasingly synthesizes agentic tasks for post-training tool-using LLMs, yet robust generalization under shifts in tasks and toolsets remains an open challenge. We trace this brittleness to insufficient diversity in synthesized training tasks. Scaling diversity is difficult because training requires tasks to remain executable and verifiable, while generalization demands diverse tool types, toolset combinations, and heterogeneous tool-use patterns. We propose DIVE, an evidence-driven recipe that inverts synthesis order, executing diverse real-world tools first and reverse-deriving tasks strictly entailed by the resulting traces, providing grounding by construction. DIVE scales structural diversity along two controllable axes, tool-pool coverage and per-task toolset variety, synthesizing 48k trajectories over 374 tools across five domains that cover 46,398 unique toolsets and 39,810 unique tool-call graphs. Training Qwen3-8B on DIVE data (48k SFT + 3.2k RL) improves by +22 average points across 9 OOD benchmarks and outperforms the strongest 8B baseline by +68%. Under a fixed budget, controlled scaling shows diversity scaling consistently outperforms quantity scaling, even with 4$\times$ less data.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26gj, title = {{DIVE}: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use}, author = {Chen, Aili and Zhang, Chi and Liu, Junteng and Chen, Jiangjie and Du, Chengyu and Li, Yunji and Zhong, Ming and Wang, Qin and Zhu, Zhengmao and Song, Jiayuan and Ji, Ke and He, Junxian and Zhao, Pengyu and Xiao, Yanghua}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {18235--18258}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26gj/chen26gj.pdf}, url = {https://proceedings.mlr.press/v306/chen26gj.html}, abstract = {Recent work increasingly synthesizes agentic tasks for post-training tool-using LLMs, yet robust generalization under shifts in tasks and toolsets remains an open challenge. We trace this brittleness to insufficient diversity in synthesized training tasks. Scaling diversity is difficult because training requires tasks to remain executable and verifiable, while generalization demands diverse tool types, toolset combinations, and heterogeneous tool-use patterns. We propose DIVE, an evidence-driven recipe that inverts synthesis order, executing diverse real-world tools first and reverse-deriving tasks strictly entailed by the resulting traces, providing grounding by construction. DIVE scales structural diversity along two controllable axes, tool-pool coverage and per-task toolset variety, synthesizing 48k trajectories over 374 tools across five domains that cover 46,398 unique toolsets and 39,810 unique tool-call graphs. Training Qwen3-8B on DIVE data (48k SFT + 3.2k RL) improves by +22 average points across 9 OOD benchmarks and outperforms the strongest 8B baseline by +68%. Under a fixed budget, controlled scaling shows diversity scaling consistently outperforms quantity scaling, even with 4$\times$ less data.} }
Endnote
%0 Conference Paper %T DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use %A Aili Chen %A Chi Zhang %A Junteng Liu %A Jiangjie Chen %A Chengyu Du %A Yunji Li %A Ming Zhong %A Qin Wang %A Zhengmao Zhu %A Jiayuan Song %A Ke Ji %A Junxian He %A Pengyu Zhao %A Yanghua Xiao %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26gj %I PMLR %P 18235--18258 %U https://proceedings.mlr.press/v306/chen26gj.html %V 306 %X Recent work increasingly synthesizes agentic tasks for post-training tool-using LLMs, yet robust generalization under shifts in tasks and toolsets remains an open challenge. We trace this brittleness to insufficient diversity in synthesized training tasks. Scaling diversity is difficult because training requires tasks to remain executable and verifiable, while generalization demands diverse tool types, toolset combinations, and heterogeneous tool-use patterns. We propose DIVE, an evidence-driven recipe that inverts synthesis order, executing diverse real-world tools first and reverse-deriving tasks strictly entailed by the resulting traces, providing grounding by construction. DIVE scales structural diversity along two controllable axes, tool-pool coverage and per-task toolset variety, synthesizing 48k trajectories over 374 tools across five domains that cover 46,398 unique toolsets and 39,810 unique tool-call graphs. Training Qwen3-8B on DIVE data (48k SFT + 3.2k RL) improves by +22 average points across 9 OOD benchmarks and outperforms the strongest 8B baseline by +68%. Under a fixed budget, controlled scaling shows diversity scaling consistently outperforms quantity scaling, even with 4$\times$ less data.
APA
Chen, A., Zhang, C., Liu, J., Chen, J., Du, C., Li, Y., Zhong, M., Wang, Q., Zhu, Z., Song, J., Ji, K., He, J., Zhao, P. & Xiao, Y.. (2026). DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:18235-18258 Available from https://proceedings.mlr.press/v306/chen26gj.html.

Related Material