SafeLab: An Interactive High-Fidelity Benchmark for Embodied Safety in Scientific Robotics

Fengshuo Bai, Yufeng Li, Ruihai Wu, Peishuo Wang, Yuhan Wang, Bernie Hao Zhu, Yuanfei Wang, Tawei Chou, Jing Gao, Runchuan Zhu, Ying Wen, Yaodong Yang, Yuanpei Chen
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:5324-5347, 2026.

Abstract

Scientific embodied agents could automate laboratory workflows, but laboratory success is trajectory-level: an agent must remain safe throughout execution, not merely reach a final goal. In benchtop settings a small pose, force, or tilt error can cause irreversible spillage or equipment damage, yet most robot benchmarks evaluate reversible, high-tolerance manipulation and imitation-trained policies receive no recovery signal for execution drift. We introduce SafeLab, a generative simulation benchmark that couples a verified generative engine, an automated expert for teleoperation-free demonstrations, and a safety-aware RL interface, with 63 calibrated laboratory assets, 64 tasks across 9 manipulation categories, and 6,400 expert trajectories. Across state-of-the-art policies, it reveals unsafe-but-successful trajectories—large gaps between task success and Safe Success Rate—most acutely in liquid handling, force-limited actuation, and bimanual glassware rearrangement. Bounded residual RL then learns execution-level corrections that raise the simulated Safe Success Rate by up to 43.0 percentage points without retraining the base policy. Finally, 50 open-loop physical replays show 86% agreement between simulated and observed safety outcomes, supporting SafeLab as a scalable platform for screening and training safer laboratory agents.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bai26e, title = {{S}afe{L}ab: An Interactive High-Fidelity Benchmark for Embodied Safety in Scientific Robotics}, author = {Bai, Fengshuo and Li, Yufeng and Wu, Ruihai and Wang, Peishuo and Wang, Yuhan and Zhu, Bernie Hao and Wang, Yuanfei and Chou, Tawei and Gao, Jing and Zhu, Runchuan and Wen, Ying and Yang, Yaodong and Chen, Yuanpei}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {5324--5347}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bai26e/bai26e.pdf}, url = {https://proceedings.mlr.press/v306/bai26e.html}, abstract = {Scientific embodied agents could automate laboratory workflows, but laboratory success is trajectory-level: an agent must remain safe throughout execution, not merely reach a final goal. In benchtop settings a small pose, force, or tilt error can cause irreversible spillage or equipment damage, yet most robot benchmarks evaluate reversible, high-tolerance manipulation and imitation-trained policies receive no recovery signal for execution drift. We introduce SafeLab, a generative simulation benchmark that couples a verified generative engine, an automated expert for teleoperation-free demonstrations, and a safety-aware RL interface, with 63 calibrated laboratory assets, 64 tasks across 9 manipulation categories, and 6,400 expert trajectories. Across state-of-the-art policies, it reveals unsafe-but-successful trajectories—large gaps between task success and Safe Success Rate—most acutely in liquid handling, force-limited actuation, and bimanual glassware rearrangement. Bounded residual RL then learns execution-level corrections that raise the simulated Safe Success Rate by up to 43.0 percentage points without retraining the base policy. Finally, 50 open-loop physical replays show 86% agreement between simulated and observed safety outcomes, supporting SafeLab as a scalable platform for screening and training safer laboratory agents.} }
Endnote
%0 Conference Paper %T SafeLab: An Interactive High-Fidelity Benchmark for Embodied Safety in Scientific Robotics %A Fengshuo Bai %A Yufeng Li %A Ruihai Wu %A Peishuo Wang %A Yuhan Wang %A Bernie Hao Zhu %A Yuanfei Wang %A Tawei Chou %A Jing Gao %A Runchuan Zhu %A Ying Wen %A Yaodong Yang %A Yuanpei Chen %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bai26e %I PMLR %P 5324--5347 %U https://proceedings.mlr.press/v306/bai26e.html %V 306 %X Scientific embodied agents could automate laboratory workflows, but laboratory success is trajectory-level: an agent must remain safe throughout execution, not merely reach a final goal. In benchtop settings a small pose, force, or tilt error can cause irreversible spillage or equipment damage, yet most robot benchmarks evaluate reversible, high-tolerance manipulation and imitation-trained policies receive no recovery signal for execution drift. We introduce SafeLab, a generative simulation benchmark that couples a verified generative engine, an automated expert for teleoperation-free demonstrations, and a safety-aware RL interface, with 63 calibrated laboratory assets, 64 tasks across 9 manipulation categories, and 6,400 expert trajectories. Across state-of-the-art policies, it reveals unsafe-but-successful trajectories—large gaps between task success and Safe Success Rate—most acutely in liquid handling, force-limited actuation, and bimanual glassware rearrangement. Bounded residual RL then learns execution-level corrections that raise the simulated Safe Success Rate by up to 43.0 percentage points without retraining the base policy. Finally, 50 open-loop physical replays show 86% agreement between simulated and observed safety outcomes, supporting SafeLab as a scalable platform for screening and training safer laboratory agents.
APA
Bai, F., Li, Y., Wu, R., Wang, P., Wang, Y., Zhu, B.H., Wang, Y., Chou, T., Gao, J., Zhu, R., Wen, Y., Yang, Y. & Chen, Y.. (2026). SafeLab: An Interactive High-Fidelity Benchmark for Embodied Safety in Scientific Robotics. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:5324-5347 Available from https://proceedings.mlr.press/v306/bai26e.html.

Related Material