InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation

Qiaosheng Chen, Yang Liu, Lei Li, Kai Chen, Qipeng Guo, Gong Cheng, Fei Yuan
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:15585-15613, 2026.

Abstract

While Large Language Models (LLMs) hold promise for automating science and education, generating interactive scientific demonstrations demands a complex synthesis of deep domain knowledge and precise reactive coding. Current benchmarks fail to capture this synergy, largely bifurcating into static code generation or text-only reasoning. To address this, we introduce InteractScience, the first benchmark dedicated to evaluating the holistic creation of interactive scientific applications. We propose a novel hybrid framework that integrates programmatic functional testing for logic verification with visually-grounded qualitative assessment for rendering fidelity. Our evaluation of 30 leading models across five disciplines reveals critical gaps in grounding scientific reasoning within interactive interfaces. By standardizing this combined capability, InteractScience establishes a crucial foundation for reliable AI-driven tools in science and education.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26cf, title = {{I}nteract{S}cience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation}, author = {Chen, Qiaosheng and Liu, Yang and Li, Lei and Chen, Kai and Guo, Qipeng and Cheng, Gong and Yuan, Fei}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {15585--15613}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26cf/chen26cf.pdf}, url = {https://proceedings.mlr.press/v306/chen26cf.html}, abstract = {While Large Language Models (LLMs) hold promise for automating science and education, generating interactive scientific demonstrations demands a complex synthesis of deep domain knowledge and precise reactive coding. Current benchmarks fail to capture this synergy, largely bifurcating into static code generation or text-only reasoning. To address this, we introduce InteractScience, the first benchmark dedicated to evaluating the holistic creation of interactive scientific applications. We propose a novel hybrid framework that integrates programmatic functional testing for logic verification with visually-grounded qualitative assessment for rendering fidelity. Our evaluation of 30 leading models across five disciplines reveals critical gaps in grounding scientific reasoning within interactive interfaces. By standardizing this combined capability, InteractScience establishes a crucial foundation for reliable AI-driven tools in science and education.} }
Endnote
%0 Conference Paper %T InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation %A Qiaosheng Chen %A Yang Liu %A Lei Li %A Kai Chen %A Qipeng Guo %A Gong Cheng %A Fei Yuan %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26cf %I PMLR %P 15585--15613 %U https://proceedings.mlr.press/v306/chen26cf.html %V 306 %X While Large Language Models (LLMs) hold promise for automating science and education, generating interactive scientific demonstrations demands a complex synthesis of deep domain knowledge and precise reactive coding. Current benchmarks fail to capture this synergy, largely bifurcating into static code generation or text-only reasoning. To address this, we introduce InteractScience, the first benchmark dedicated to evaluating the holistic creation of interactive scientific applications. We propose a novel hybrid framework that integrates programmatic functional testing for logic verification with visually-grounded qualitative assessment for rendering fidelity. Our evaluation of 30 leading models across five disciplines reveals critical gaps in grounding scientific reasoning within interactive interfaces. By standardizing this combined capability, InteractScience establishes a crucial foundation for reliable AI-driven tools in science and education.
APA
Chen, Q., Liu, Y., Li, L., Chen, K., Guo, Q., Cheng, G. & Yuan, F.. (2026). InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:15585-15613 Available from https://proceedings.mlr.press/v306/chen26cf.html.

Related Material