ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research

Ruicheng Ao, David Simchi-Levi, Xinshang Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:3056-3102, 2026.

Abstract

Operations Research practitioners debug infeasible models through an iterative process: inspecting Irreducible Infeasible Subsystems (IIS), identifying constraint conflicts, and repairing formulations until feasibility is restored. Existing LLM benchmarks mostly treat OR as one-shot translation from problem descriptions to solver code, omitting this diagnostic loop. We formalize infeasible-model repair as a solver-in-the-loop Markov Decision Process in which each action triggers solver re-execution and IIS recomputation, yielding deterministic, verifiable feedback. We introduce ORLoopBench, a benchmark suite with two components: ORDebug releases 5,362 LP/MILP repair instances, while ORBias evaluates closed-form operational decision rationality across inventory settings. Solver-verified RLVR training enables an 8B model to surpass frontier APIs on LP repair (95.3% vs 92.4% RR@5), improves diagnostic behavior, and transfers to MILP repair. The same evaluation exposes semantic drift in whole-model code regeneration: feasible regenerated MILPs can solve the wrong problem. Process-level evaluation with solver oracles enables targeted training for reliable OR self-correction.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-ao26a, title = {{ORL}oop{B}ench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research}, author = {Ao, Ruicheng and Simchi-Levi, David and Wang, Xinshang}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {3056--3102}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ao26a/ao26a.pdf}, url = {https://proceedings.mlr.press/v306/ao26a.html}, abstract = {Operations Research practitioners debug infeasible models through an iterative process: inspecting Irreducible Infeasible Subsystems (IIS), identifying constraint conflicts, and repairing formulations until feasibility is restored. Existing LLM benchmarks mostly treat OR as one-shot translation from problem descriptions to solver code, omitting this diagnostic loop. We formalize infeasible-model repair as a solver-in-the-loop Markov Decision Process in which each action triggers solver re-execution and IIS recomputation, yielding deterministic, verifiable feedback. We introduce ORLoopBench, a benchmark suite with two components: ORDebug releases 5,362 LP/MILP repair instances, while ORBias evaluates closed-form operational decision rationality across inventory settings. Solver-verified RLVR training enables an 8B model to surpass frontier APIs on LP repair (95.3% vs 92.4% RR@5), improves diagnostic behavior, and transfers to MILP repair. The same evaluation exposes semantic drift in whole-model code regeneration: feasible regenerated MILPs can solve the wrong problem. Process-level evaluation with solver oracles enables targeted training for reliable OR self-correction.} }
Endnote
%0 Conference Paper %T ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research %A Ruicheng Ao %A David Simchi-Levi %A Xinshang Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-ao26a %I PMLR %P 3056--3102 %U https://proceedings.mlr.press/v306/ao26a.html %V 306 %X Operations Research practitioners debug infeasible models through an iterative process: inspecting Irreducible Infeasible Subsystems (IIS), identifying constraint conflicts, and repairing formulations until feasibility is restored. Existing LLM benchmarks mostly treat OR as one-shot translation from problem descriptions to solver code, omitting this diagnostic loop. We formalize infeasible-model repair as a solver-in-the-loop Markov Decision Process in which each action triggers solver re-execution and IIS recomputation, yielding deterministic, verifiable feedback. We introduce ORLoopBench, a benchmark suite with two components: ORDebug releases 5,362 LP/MILP repair instances, while ORBias evaluates closed-form operational decision rationality across inventory settings. Solver-verified RLVR training enables an 8B model to surpass frontier APIs on LP repair (95.3% vs 92.4% RR@5), improves diagnostic behavior, and transfers to MILP repair. The same evaluation exposes semantic drift in whole-model code regeneration: feasible regenerated MILPs can solve the wrong problem. Process-level evaluation with solver oracles enables targeted training for reliable OR self-correction.
APA
Ao, R., Simchi-Levi, D. & Wang, X.. (2026). ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:3056-3102 Available from https://proceedings.mlr.press/v306/ao26a.html.

Related Material