Optimal Query Allocation in Extractive QA with LLMs: A Learning-to-Defer Framework with Theoretical Guarantees

Yannis Montreuil, Yeo Shu Heng, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:2665-2673, 2026.

Abstract

Large Language Models (LLMs) excel at generative language tasks but remain unreliable for structured prediction, particularly in extractive question answering (EQA), where success depends on precise span selection. These challenges are amplified in resource-constrained environments, such as mobile or embedded systems, where deploying high-capacity models is often infeasible. We propose a Learning-to-Defer framework that routes EQA queries across a pool of models with varying capabilities and costs to balance accuracy and efficiency. Our approach is grounded in statistical decision theory: we define a differentiable surrogate loss whose minimizer provably converges to the Bayes-optimal allocation policy. Experiments on SQuADv1, SQuADv2, and TriviaQA show that our method consistently improves the accuracy-efficiency trade-off relative to static baselines and prior routing heuristics. Overall, our framework provides a principled and scalable solution for EQA in both high-performance and on-device deployment settings.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-montreuil26a, title = { Optimal Query Allocation in Extractive QA with LLMs: A Learning-to-Defer Framework with Theoretical Guarantees }, author = {Montreuil, Yannis and Heng, Yeo Shu and Carlier, Axel and Ng, Lai Xing and Ooi, Wei Tsang}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {2665--2673}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/montreuil26a/montreuil26a.pdf}, url = {https://proceedings.mlr.press/v300/montreuil26a.html}, abstract = { Large Language Models (LLMs) excel at generative language tasks but remain unreliable for structured prediction, particularly in extractive question answering (EQA), where success depends on precise span selection. These challenges are amplified in resource-constrained environments, such as mobile or embedded systems, where deploying high-capacity models is often infeasible. We propose a Learning-to-Defer framework that routes EQA queries across a pool of models with varying capabilities and costs to balance accuracy and efficiency. Our approach is grounded in statistical decision theory: we define a differentiable surrogate loss whose minimizer provably converges to the Bayes-optimal allocation policy. Experiments on SQuADv1, SQuADv2, and TriviaQA show that our method consistently improves the accuracy-efficiency trade-off relative to static baselines and prior routing heuristics. Overall, our framework provides a principled and scalable solution for EQA in both high-performance and on-device deployment settings. } }
Endnote
%0 Conference Paper %T Optimal Query Allocation in Extractive QA with LLMs: A Learning-to-Defer Framework with Theoretical Guarantees %A Yannis Montreuil %A Yeo Shu Heng %A Axel Carlier %A Lai Xing Ng %A Wei Tsang Ooi %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-montreuil26a %I PMLR %P 2665--2673 %U https://proceedings.mlr.press/v300/montreuil26a.html %V 300 %X Large Language Models (LLMs) excel at generative language tasks but remain unreliable for structured prediction, particularly in extractive question answering (EQA), where success depends on precise span selection. These challenges are amplified in resource-constrained environments, such as mobile or embedded systems, where deploying high-capacity models is often infeasible. We propose a Learning-to-Defer framework that routes EQA queries across a pool of models with varying capabilities and costs to balance accuracy and efficiency. Our approach is grounded in statistical decision theory: we define a differentiable surrogate loss whose minimizer provably converges to the Bayes-optimal allocation policy. Experiments on SQuADv1, SQuADv2, and TriviaQA show that our method consistently improves the accuracy-efficiency trade-off relative to static baselines and prior routing heuristics. Overall, our framework provides a principled and scalable solution for EQA in both high-performance and on-device deployment settings.
APA
Montreuil, Y., Heng, Y.S., Carlier, A., Ng, L.X. & Ooi, W.T.. (2026). Optimal Query Allocation in Extractive QA with LLMs: A Learning-to-Defer Framework with Theoretical Guarantees . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:2665-2673 Available from https://proceedings.mlr.press/v300/montreuil26a.html.

Related Material