IDK Cascades: Fast Deep Learning by Learning not to Overthink

Xin Wang, Yujia Luo, Daniel Crankshaw, Alexey Tumanov, Fisher Yu, Joseph E. Gonzalez
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:579-589, 2018.

Abstract

Advances in deep learning have led to substan- tial increases in prediction accuracy but have been accompanied by increases in the cost of rendering predictions. We conjecture that for a majority of real-world inputs, the recent ad- vances in deep learning have created models that effectively “over-think” on simple inputs. In this paper we revisit the classic question of building model cascades that primarily leverage class asymmetry to reduce cost. We introduce the “I Don’t Know” (IDK) prediction cascades framework, a general framework to systemat- ically compose a set of pre-trained models to accelerate inference without a loss in predic- tion accuracy. We propose two search based methods for constructing cascades as well as a new cost-aware objective within this frame- work. The proposed IDK cascade framework can be easily adopted in the existing model serving systems without additional model re- training. We evaluate the proposed techniques on a range of benchmarks to demonstrate the effectiveness of the proposed framework.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-wang18c, title = {{IDK} Cascades: Fast Deep Learning by Learning not to Overthink}, author = {Wang, Xin and Luo, Yujia and Crankshaw, Daniel and Tumanov, Alexey and Yu, Fisher and Gonzalez, Joseph E.}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {579--589}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/wang18c/wang18c.pdf}, url = {https://proceedings.mlr.press/r16/wang18c.html}, abstract = {Advances in deep learning have led to substan- tial increases in prediction accuracy but have been accompanied by increases in the cost of rendering predictions. We conjecture that for a majority of real-world inputs, the recent ad- vances in deep learning have created models that effectively “over-think” on simple inputs. In this paper we revisit the classic question of building model cascades that primarily leverage class asymmetry to reduce cost. We introduce the “I Don’t Know” (IDK) prediction cascades framework, a general framework to systemat- ically compose a set of pre-trained models to accelerate inference without a loss in predic- tion accuracy. We propose two search based methods for constructing cascades as well as a new cost-aware objective within this frame- work. The proposed IDK cascade framework can be easily adopted in the existing model serving systems without additional model re- training. We evaluate the proposed techniques on a range of benchmarks to demonstrate the effectiveness of the proposed framework.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T IDK Cascades: Fast Deep Learning by Learning not to Overthink %A Xin Wang %A Yujia Luo %A Daniel Crankshaw %A Alexey Tumanov %A Fisher Yu %A Joseph E. Gonzalez %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-wang18c %I PMLR %P 579--589 %U https://proceedings.mlr.press/r16/wang18c.html %V R16 %X Advances in deep learning have led to substan- tial increases in prediction accuracy but have been accompanied by increases in the cost of rendering predictions. We conjecture that for a majority of real-world inputs, the recent ad- vances in deep learning have created models that effectively “over-think” on simple inputs. In this paper we revisit the classic question of building model cascades that primarily leverage class asymmetry to reduce cost. We introduce the “I Don’t Know” (IDK) prediction cascades framework, a general framework to systemat- ically compose a set of pre-trained models to accelerate inference without a loss in predic- tion accuracy. We propose two search based methods for constructing cascades as well as a new cost-aware objective within this frame- work. The proposed IDK cascade framework can be easily adopted in the existing model serving systems without additional model re- training. We evaluate the proposed techniques on a range of benchmarks to demonstrate the effectiveness of the proposed framework. %Z Reissued by PMLR on 04 October 2026.
APA
Wang, X., Luo, Y., Crankshaw, D., Tumanov, A., Yu, F. & Gonzalez, J.E.. (2018). IDK Cascades: Fast Deep Learning by Learning not to Overthink. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:579-589 Available from https://proceedings.mlr.press/r16/wang18c.html. Reissued by PMLR on 04 October 2026.

Related Material