Diffract: Spectral View of LLM Domain Adaptation

Nikita Borodin, Maria Krylova, Artem Zabolotnyi, Dmitry Aspisov, Egor Shikov, Nikita Tyuplyaev, Oleg Travkin, Roman Alferov, Dmitry Vinichenko
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:9269-9294, 2026.

Abstract

We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text. Using singular value decomposition of weight matrices, we find that CPT leaves singular value spectra largely invariant, with adaptation driven mainly by changes in singular vectors. An analysis of attention-head projection matrices reveals strong, domain-dependent head heterogeneity, which we exploit to define a head importance criterion: up to 60% of head updates can be removed without measurable quality loss. Selectively rewinding low-importance heads to their pre-trained state improves benchmark accuracy by up to 4% versus the fully trained baseline. Finally, we identify domain connectivity—linear interpolation between CPT checkpoints yields smooth domain-quality interpolation without notable degradation on either domain—and release Diffract, an open-source toolkit for scalable spectral analysis of billion-parameter models.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-borodin26a, title = {Diffract: Spectral View of {LLM} Domain Adaptation}, author = {Borodin, Nikita and Krylova, Maria and Zabolotnyi, Artem and Aspisov, Dmitry and Shikov, Egor and Tyuplyaev, Nikita and Travkin, Oleg and Alferov, Roman and Vinichenko, Dmitry}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {9269--9294}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/borodin26a/borodin26a.pdf}, url = {https://proceedings.mlr.press/v306/borodin26a.html}, abstract = {We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text. Using singular value decomposition of weight matrices, we find that CPT leaves singular value spectra largely invariant, with adaptation driven mainly by changes in singular vectors. An analysis of attention-head projection matrices reveals strong, domain-dependent head heterogeneity, which we exploit to define a head importance criterion: up to 60% of head updates can be removed without measurable quality loss. Selectively rewinding low-importance heads to their pre-trained state improves benchmark accuracy by up to 4% versus the fully trained baseline. Finally, we identify domain connectivity—linear interpolation between CPT checkpoints yields smooth domain-quality interpolation without notable degradation on either domain—and release Diffract, an open-source toolkit for scalable spectral analysis of billion-parameter models.} }
Endnote
%0 Conference Paper %T Diffract: Spectral View of LLM Domain Adaptation %A Nikita Borodin %A Maria Krylova %A Artem Zabolotnyi %A Dmitry Aspisov %A Egor Shikov %A Nikita Tyuplyaev %A Oleg Travkin %A Roman Alferov %A Dmitry Vinichenko %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-borodin26a %I PMLR %P 9269--9294 %U https://proceedings.mlr.press/v306/borodin26a.html %V 306 %X We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text. Using singular value decomposition of weight matrices, we find that CPT leaves singular value spectra largely invariant, with adaptation driven mainly by changes in singular vectors. An analysis of attention-head projection matrices reveals strong, domain-dependent head heterogeneity, which we exploit to define a head importance criterion: up to 60% of head updates can be removed without measurable quality loss. Selectively rewinding low-importance heads to their pre-trained state improves benchmark accuracy by up to 4% versus the fully trained baseline. Finally, we identify domain connectivity—linear interpolation between CPT checkpoints yields smooth domain-quality interpolation without notable degradation on either domain—and release Diffract, an open-source toolkit for scalable spectral analysis of billion-parameter models.
APA
Borodin, N., Krylova, M., Zabolotnyi, A., Aspisov, D., Shikov, E., Tyuplyaev, N., Travkin, O., Alferov, R. & Vinichenko, D.. (2026). Diffract: Spectral View of LLM Domain Adaptation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:9269-9294 Available from https://proceedings.mlr.press/v306/borodin26a.html.

Related Material