Unifying Stacking and Cascading for Efficient Ensemble Inference

Ashwin Gerard Colaco, Sharad Mehrotra, Michael J. De Lucia, Kevin Hamlen, Murat Kantarcioglu, Latifur Khan, Ananthram Swami, Bhavani Thuraisingham, Unnat Jain
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:21176-21211, 2026.

Abstract

We introduce LazyStack, a method for efficient model ensemble inference. The core idea is intuitive: after each model executes, we check whether accumulated evidence is sufficient to exit confidently. Sometimes one model suffices; other times we aggregate predictions from several models via trained meta-learners before reaching confidence. Two insights make this work. First, most inputs follow only 3 to 8 execution trajectories. This reduces the training problem from exponential to linear: we learn aggregators only for these common paths, not all possible model combinations. Second, we formulate trajectory selection as an MDP and use value iteration to compute the optimal routing policy, which reveals counterintuitive model orderings. On intrusion detection, starting with a moderately expensive model outperforms starting with the cheapest, because its higher confidence enables earlier overall exit. Across vision, text, tabular, and LLM tasks, we achieve up to 38x speedup at 97%+ accuracy retention compared to a complete ensemble. The result: ensemble-quality predictions at cascade-level cost. Code and a project page are available at https://ashwincolaco.github.io/lazystack.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-colaco26a, title = {Unifying Stacking and Cascading for Efficient Ensemble Inference}, author = {Colaco, Ashwin Gerard and Mehrotra, Sharad and De Lucia, Michael J. and Hamlen, Kevin and Kantarcioglu, Murat and Khan, Latifur and Swami, Ananthram and Thuraisingham, Bhavani and Jain, Unnat}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {21176--21211}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/colaco26a/colaco26a.pdf}, url = {https://proceedings.mlr.press/v306/colaco26a.html}, abstract = {We introduce LazyStack, a method for efficient model ensemble inference. The core idea is intuitive: after each model executes, we check whether accumulated evidence is sufficient to exit confidently. Sometimes one model suffices; other times we aggregate predictions from several models via trained meta-learners before reaching confidence. Two insights make this work. First, most inputs follow only 3 to 8 execution trajectories. This reduces the training problem from exponential to linear: we learn aggregators only for these common paths, not all possible model combinations. Second, we formulate trajectory selection as an MDP and use value iteration to compute the optimal routing policy, which reveals counterintuitive model orderings. On intrusion detection, starting with a moderately expensive model outperforms starting with the cheapest, because its higher confidence enables earlier overall exit. Across vision, text, tabular, and LLM tasks, we achieve up to 38x speedup at 97%+ accuracy retention compared to a complete ensemble. The result: ensemble-quality predictions at cascade-level cost. Code and a project page are available at https://ashwincolaco.github.io/lazystack.} }
Endnote
%0 Conference Paper %T Unifying Stacking and Cascading for Efficient Ensemble Inference %A Ashwin Gerard Colaco %A Sharad Mehrotra %A Michael J. De Lucia %A Kevin Hamlen %A Murat Kantarcioglu %A Latifur Khan %A Ananthram Swami %A Bhavani Thuraisingham %A Unnat Jain %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-colaco26a %I PMLR %P 21176--21211 %U https://proceedings.mlr.press/v306/colaco26a.html %V 306 %X We introduce LazyStack, a method for efficient model ensemble inference. The core idea is intuitive: after each model executes, we check whether accumulated evidence is sufficient to exit confidently. Sometimes one model suffices; other times we aggregate predictions from several models via trained meta-learners before reaching confidence. Two insights make this work. First, most inputs follow only 3 to 8 execution trajectories. This reduces the training problem from exponential to linear: we learn aggregators only for these common paths, not all possible model combinations. Second, we formulate trajectory selection as an MDP and use value iteration to compute the optimal routing policy, which reveals counterintuitive model orderings. On intrusion detection, starting with a moderately expensive model outperforms starting with the cheapest, because its higher confidence enables earlier overall exit. Across vision, text, tabular, and LLM tasks, we achieve up to 38x speedup at 97%+ accuracy retention compared to a complete ensemble. The result: ensemble-quality predictions at cascade-level cost. Code and a project page are available at https://ashwincolaco.github.io/lazystack.
APA
Colaco, A.G., Mehrotra, S., De Lucia, M.J., Hamlen, K., Kantarcioglu, M., Khan, L., Swami, A., Thuraisingham, B. & Jain, U.. (2026). Unifying Stacking and Cascading for Efficient Ensemble Inference. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:21176-21211 Available from https://proceedings.mlr.press/v306/colaco26a.html.

Related Material