Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing

Raghavv Goel, Mukul Gagrani, Mingu Lee, Christopher Lott
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:35269-35290, 2026.

Abstract

Large Language Models (LLMs) possess latent multi-token prediction (MTP) abilities despite being trained only for next-token generation. We introduce ESP (Embedding-Space Probing), a simple and training-free MTP method that probes an LLM using on-the-fly mask tokens drawn from its embedding space, enabling parallel future-token prediction without modifying weights or relying on draft models. ESP constructs a speculative token tree by sampling Top-K candidates from mask-token logits and applies a lightweight pruning rule to retain high-probability continuations. During generation, predictions are verified in parallel, yielding lossless decoding while significantly reducing model calls and increasing token throughput. ESP consistently outperforms existing training-free baselines, improving acceptance length by $7$–$11$% over Lookahead Decoding on LLaMA3 and $7$–$8$% on Qwen3, and increasing throughput by up to $15$–$19$% over the strongest baseline. Finally, we provide theoretical insight and empirical evidence showing that decoder layers naturally align mask-token representations with next-token states, enabling accurate multi-step prediction without retraining or auxiliary models.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-goel26a, title = {Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing}, author = {Goel, Raghavv and Gagrani, Mukul and Lee, Mingu and Lott, Christopher}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {35269--35290}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/goel26a/goel26a.pdf}, url = {https://proceedings.mlr.press/v306/goel26a.html}, abstract = {Large Language Models (LLMs) possess latent multi-token prediction (MTP) abilities despite being trained only for next-token generation. We introduce ESP (Embedding-Space Probing), a simple and training-free MTP method that probes an LLM using on-the-fly mask tokens drawn from its embedding space, enabling parallel future-token prediction without modifying weights or relying on draft models. ESP constructs a speculative token tree by sampling Top-K candidates from mask-token logits and applies a lightweight pruning rule to retain high-probability continuations. During generation, predictions are verified in parallel, yielding lossless decoding while significantly reducing model calls and increasing token throughput. ESP consistently outperforms existing training-free baselines, improving acceptance length by $7$–$11$% over Lookahead Decoding on LLaMA3 and $7$–$8$% on Qwen3, and increasing throughput by up to $15$–$19$% over the strongest baseline. Finally, we provide theoretical insight and empirical evidence showing that decoder layers naturally align mask-token representations with next-token states, enabling accurate multi-step prediction without retraining or auxiliary models.} }
Endnote
%0 Conference Paper %T Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing %A Raghavv Goel %A Mukul Gagrani %A Mingu Lee %A Christopher Lott %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-goel26a %I PMLR %P 35269--35290 %U https://proceedings.mlr.press/v306/goel26a.html %V 306 %X Large Language Models (LLMs) possess latent multi-token prediction (MTP) abilities despite being trained only for next-token generation. We introduce ESP (Embedding-Space Probing), a simple and training-free MTP method that probes an LLM using on-the-fly mask tokens drawn from its embedding space, enabling parallel future-token prediction without modifying weights or relying on draft models. ESP constructs a speculative token tree by sampling Top-K candidates from mask-token logits and applies a lightweight pruning rule to retain high-probability continuations. During generation, predictions are verified in parallel, yielding lossless decoding while significantly reducing model calls and increasing token throughput. ESP consistently outperforms existing training-free baselines, improving acceptance length by $7$–$11$% over Lookahead Decoding on LLaMA3 and $7$–$8$% on Qwen3, and increasing throughput by up to $15$–$19$% over the strongest baseline. Finally, we provide theoretical insight and empirical evidence showing that decoder layers naturally align mask-token representations with next-token states, enabling accurate multi-step prediction without retraining or auxiliary models.
APA
Goel, R., Gagrani, M., Lee, M. & Lott, C.. (2026). Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:35269-35290 Available from https://proceedings.mlr.press/v306/goel26a.html.

Related Material