The Bridge-Garden Dilemma in LLM Distillation: Why Mixing Hard and Soft Labels Works

Guanghui Wang, Kaiwen Lv Kacuila, Zhiyong Yang, Zitai Wang, Jin-Wen Wu, Longtao Huang, Qianqian Xu, Qingming Huang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:127315-127354, 2026.

Abstract

Knowledge distillation (KD) transfers knowledge from a large teacher model to a smaller student. In language modeling, the student is trained either on tokens sampled from the teacher (hard labels) or the teacher’s full next-token distribution (soft labels). Despite soft labels appear strictly richer, we find that mixing hard and soft labels consistently yields better results. Crucially, we show that this gain cannot be explained by closer teacher matching during training. Instead, it comes from reduced exposure bias—the mismatch between training and inference distributions. To explain this phenomenon, we introduce the Bridge–Garden Decomposition theory, which categorizes generation steps into two types: Bridges, where the next token must be exact, and Gardens, where it can be flexible. We show that hard-only KD excels in Bridges by avoiding risky deviations, while soft-only KD preserves diversity in Gardens. A hybrid strategy handles both cases and, as a result, reduces exposure bias across the sequence. Guided by this theory, we develop a family of Bridge–Garden hybrid supervision methods that adaptively balance hard and soft labels. Across seven teacher–student pairs (including Qwen, Llama, Gemma, and DeepSeek) and benchmarks in reasoning and coding, our approach outperforms divergence-based and on-policy KD baselines while reducing training cost by 9.7$\times$, enabling efficient model compression.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-wang26cw, title = {The Bridge-Garden Dilemma in {LLM} Distillation: Why Mixing Hard and Soft Labels Works}, author = {Wang, Guanghui and Kacuila, Kaiwen Lv and Yang, Zhiyong and Wang, Zitai and Wu, Jin-Wen and Huang, Longtao and Xu, Qianqian and Huang, Qingming}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {127315--127354}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/wang26cw/wang26cw.pdf}, url = {https://proceedings.mlr.press/v306/wang26cw.html}, abstract = {Knowledge distillation (KD) transfers knowledge from a large teacher model to a smaller student. In language modeling, the student is trained either on tokens sampled from the teacher (hard labels) or the teacher’s full next-token distribution (soft labels). Despite soft labels appear strictly richer, we find that mixing hard and soft labels consistently yields better results. Crucially, we show that this gain cannot be explained by closer teacher matching during training. Instead, it comes from reduced exposure bias—the mismatch between training and inference distributions. To explain this phenomenon, we introduce the Bridge–Garden Decomposition theory, which categorizes generation steps into two types: Bridges, where the next token must be exact, and Gardens, where it can be flexible. We show that hard-only KD excels in Bridges by avoiding risky deviations, while soft-only KD preserves diversity in Gardens. A hybrid strategy handles both cases and, as a result, reduces exposure bias across the sequence. Guided by this theory, we develop a family of Bridge–Garden hybrid supervision methods that adaptively balance hard and soft labels. Across seven teacher–student pairs (including Qwen, Llama, Gemma, and DeepSeek) and benchmarks in reasoning and coding, our approach outperforms divergence-based and on-policy KD baselines while reducing training cost by 9.7$\times$, enabling efficient model compression.} }
Endnote
%0 Conference Paper %T The Bridge-Garden Dilemma in LLM Distillation: Why Mixing Hard and Soft Labels Works %A Guanghui Wang %A Kaiwen Lv Kacuila %A Zhiyong Yang %A Zitai Wang %A Jin-Wen Wu %A Longtao Huang %A Qianqian Xu %A Qingming Huang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-wang26cw %I PMLR %P 127315--127354 %U https://proceedings.mlr.press/v306/wang26cw.html %V 306 %X Knowledge distillation (KD) transfers knowledge from a large teacher model to a smaller student. In language modeling, the student is trained either on tokens sampled from the teacher (hard labels) or the teacher’s full next-token distribution (soft labels). Despite soft labels appear strictly richer, we find that mixing hard and soft labels consistently yields better results. Crucially, we show that this gain cannot be explained by closer teacher matching during training. Instead, it comes from reduced exposure bias—the mismatch between training and inference distributions. To explain this phenomenon, we introduce the Bridge–Garden Decomposition theory, which categorizes generation steps into two types: Bridges, where the next token must be exact, and Gardens, where it can be flexible. We show that hard-only KD excels in Bridges by avoiding risky deviations, while soft-only KD preserves diversity in Gardens. A hybrid strategy handles both cases and, as a result, reduces exposure bias across the sequence. Guided by this theory, we develop a family of Bridge–Garden hybrid supervision methods that adaptively balance hard and soft labels. Across seven teacher–student pairs (including Qwen, Llama, Gemma, and DeepSeek) and benchmarks in reasoning and coding, our approach outperforms divergence-based and on-policy KD baselines while reducing training cost by 9.7$\times$, enabling efficient model compression.
APA
Wang, G., Kacuila, K.L., Yang, Z., Wang, Z., Wu, J., Huang, L., Xu, Q. & Huang, Q.. (2026). The Bridge-Garden Dilemma in LLM Distillation: Why Mixing Hard and Soft Labels Works. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:127315-127354 Available from https://proceedings.mlr.press/v306/wang26cw.html.

Related Material