LENS: Latent Precision Inference in Multi-LLM Routing

Juntao Liu, Lixing Yu, Kun Yue, Zhiwen Tang
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3933-3956, 2026.

Abstract

Large language model ({LLM}) routing aims to select an appropriate model for each query under performance–cost trade-offs. A natural approach to improve adaptivity is to leverage interaction feedback to form behavioral signatures and continuously refine routing decisions as the environment changes. However, in realistic deployments, such feedback can have highly variable effective precision due to selective logging, imperfect evaluation signals, and temporal drift. As a result, behavioral signatures may be noisy or weakly informative, and treating them as uniformly precise can miscalibrate performance estimation and destabilize routing effectiveness. We propose the \textbf{L}atent pr\textbf{E}cisio\textbf{N} inference \textbf{S}ystem (\textbf{LENS}), a probabilistic routing framework that explicitly models the latent precision of interaction-derived signals. LENS formulates routing as posterior utility maximization under imprecise supervision, and marginalizes over latent precision to adaptively control how strongly behavioral signatures influence model selection. We instantiate LENS with an efficient variational inference procedure and evaluate it on multi-{LLM} routing benchmarks across diverse tasks and distribution shifts. Experimental results show that LENS consistently improves performance–cost trade-offs, with particularly strong gains under task and model shifts.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-liu26b, title = {LENS: Latent Precision Inference in Multi-{LLM} Routing}, author = {Liu, Juntao and Yu, Lixing and Yue, Kun and Tang, Zhiwen}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {3933--3956}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/liu26b/liu26b.pdf}, url = {https://proceedings.mlr.press/v337/liu26b.html}, abstract = {Large language model ({LLM}) routing aims to select an appropriate model for each query under performance–cost trade-offs. A natural approach to improve adaptivity is to leverage interaction feedback to form behavioral signatures and continuously refine routing decisions as the environment changes. However, in realistic deployments, such feedback can have highly variable effective precision due to selective logging, imperfect evaluation signals, and temporal drift. As a result, behavioral signatures may be noisy or weakly informative, and treating them as uniformly precise can miscalibrate performance estimation and destabilize routing effectiveness. We propose the \textbf{L}atent pr\textbf{E}cisio\textbf{N} inference \textbf{S}ystem (\textbf{LENS}), a probabilistic routing framework that explicitly models the latent precision of interaction-derived signals. LENS formulates routing as posterior utility maximization under imprecise supervision, and marginalizes over latent precision to adaptively control how strongly behavioral signatures influence model selection. We instantiate LENS with an efficient variational inference procedure and evaluate it on multi-{LLM} routing benchmarks across diverse tasks and distribution shifts. Experimental results show that LENS consistently improves performance–cost trade-offs, with particularly strong gains under task and model shifts.} }
Endnote
%0 Conference Paper %T LENS: Latent Precision Inference in Multi-LLM Routing %A Juntao Liu %A Lixing Yu %A Kun Yue %A Zhiwen Tang %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-liu26b %I PMLR %P 3933--3956 %U https://proceedings.mlr.press/v337/liu26b.html %V 337 %X Large language model ({LLM}) routing aims to select an appropriate model for each query under performance–cost trade-offs. A natural approach to improve adaptivity is to leverage interaction feedback to form behavioral signatures and continuously refine routing decisions as the environment changes. However, in realistic deployments, such feedback can have highly variable effective precision due to selective logging, imperfect evaluation signals, and temporal drift. As a result, behavioral signatures may be noisy or weakly informative, and treating them as uniformly precise can miscalibrate performance estimation and destabilize routing effectiveness. We propose the \textbf{L}atent pr\textbf{E}cisio\textbf{N} inference \textbf{S}ystem (\textbf{LENS}), a probabilistic routing framework that explicitly models the latent precision of interaction-derived signals. LENS formulates routing as posterior utility maximization under imprecise supervision, and marginalizes over latent precision to adaptively control how strongly behavioral signatures influence model selection. We instantiate LENS with an efficient variational inference procedure and evaluate it on multi-{LLM} routing benchmarks across diverse tasks and distribution shifts. Experimental results show that LENS consistently improves performance–cost trade-offs, with particularly strong gains under task and model shifts.
APA
Liu, J., Yu, L., Yue, K. & Tang, Z.. (2026). LENS: Latent Precision Inference in Multi-LLM Routing. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:3933-3956 Available from https://proceedings.mlr.press/v337/liu26b.html.

Related Material