SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs

Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence, Rohan Bhakta, Izzy F. Thomas, Mingyun Cao, Chung-Hao Tsai, Zihao Zhou, Yidong Zhao, Hao Liu, Alessandro Marinoni, Alexey Arefiev, Rose Yu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:11206-11254, 2026.

Abstract

Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. SimulCost compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,643 single-round (initial guess) and 2,304 multi-round (adjustment by trial-and-error) tasks across 11 simulators from fluid dynamics, solid mechanics, and plasma physics, whose costs are analytically defined and platform-independent. A twelfth simulator, a production plasma code measurable only by wall clock, is reported separately. Frontier LLMs achieve 45–62% success rates in single-round mode, dropping to 34–50% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 66–81%, but LLMs are 1.5–2.7$\times$ slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source SimulCost as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cao26g, title = {{S}imul{C}ost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with {LLM}s}, author = {Cao, Yadi and Lai, Sicheng and Huang, Jiahe and Zhang, Yang and Lawrence, Zach and Bhakta, Rohan and Thomas, Izzy F. and Cao, Mingyun and Tsai, Chung-Hao and Zhou, Zihao and Zhao, Yidong and Liu, Hao and Marinoni, Alessandro and Arefiev, Alexey and Yu, Rose}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {11206--11254}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cao26g/cao26g.pdf}, url = {https://proceedings.mlr.press/v306/cao26g.html}, abstract = {Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. SimulCost compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,643 single-round (initial guess) and 2,304 multi-round (adjustment by trial-and-error) tasks across 11 simulators from fluid dynamics, solid mechanics, and plasma physics, whose costs are analytically defined and platform-independent. A twelfth simulator, a production plasma code measurable only by wall clock, is reported separately. Frontier LLMs achieve 45–62% success rates in single-round mode, dropping to 34–50% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 66–81%, but LLMs are 1.5–2.7$\times$ slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source SimulCost as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench} }
Endnote
%0 Conference Paper %T SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs %A Yadi Cao %A Sicheng Lai %A Jiahe Huang %A Yang Zhang %A Zach Lawrence %A Rohan Bhakta %A Izzy F. Thomas %A Mingyun Cao %A Chung-Hao Tsai %A Zihao Zhou %A Yidong Zhao %A Hao Liu %A Alessandro Marinoni %A Alexey Arefiev %A Rose Yu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cao26g %I PMLR %P 11206--11254 %U https://proceedings.mlr.press/v306/cao26g.html %V 306 %X Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. SimulCost compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,643 single-round (initial guess) and 2,304 multi-round (adjustment by trial-and-error) tasks across 11 simulators from fluid dynamics, solid mechanics, and plasma physics, whose costs are analytically defined and platform-independent. A twelfth simulator, a production plasma code measurable only by wall clock, is reported separately. Frontier LLMs achieve 45–62% success rates in single-round mode, dropping to 34–50% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 66–81%, but LLMs are 1.5–2.7$\times$ slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source SimulCost as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench
APA
Cao, Y., Lai, S., Huang, J., Zhang, Y., Lawrence, Z., Bhakta, R., Thomas, I.F., Cao, M., Tsai, C., Zhou, Z., Zhao, Y., Liu, H., Marinoni, A., Arefiev, A. & Yu, R.. (2026). SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:11206-11254 Available from https://proceedings.mlr.press/v306/cao26g.html.

Related Material