Skill Neologisms: Towards Skill-based Continual Learning

Antonin Berthon, Nicolás Astorga, Mihaela Van Der Schaar
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:7815-7836, 2026.

Abstract

Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending model capabilities to new skills in a scalable manner is an open problem: fine-tuning and parameter-efficient variants risk catastrophic forgetting, while context-based approaches have limited expressiveness and are constrained by the model’s effective context. We explore skill neologisms–soft tokens integrated in the model’s vocabulary and optimized to improve capabilities over a specific skill–as a way to selectively acquire new skills without weight updates. We first observe that pretrained LLMs already exhibit tokens associated with procedural knowledge. We then show on a controlled synthetic task that skill neologisms can be learned to improve model capabilities on specific skills while being composable with out-of-distribution skills, and that independently trained skill neologisms can be composed zero-shot. Finally, we validate zero-shot composition of independently learned skill neologisms on the more realistic natural language setting of the Skill-Mix benchmark. These results suggest that skill neologisms may provide a scalable path towards skill-based continual learning.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-berthon26a, title = {Skill Neologisms: Towards Skill-based Continual Learning}, author = {Berthon, Antonin and Astorga, Nicol\'{a}s and Van Der Schaar, Mihaela}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {7815--7836}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/berthon26a/berthon26a.pdf}, url = {https://proceedings.mlr.press/v306/berthon26a.html}, abstract = {Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending model capabilities to new skills in a scalable manner is an open problem: fine-tuning and parameter-efficient variants risk catastrophic forgetting, while context-based approaches have limited expressiveness and are constrained by the model’s effective context. We explore skill neologisms–soft tokens integrated in the model’s vocabulary and optimized to improve capabilities over a specific skill–as a way to selectively acquire new skills without weight updates. We first observe that pretrained LLMs already exhibit tokens associated with procedural knowledge. We then show on a controlled synthetic task that skill neologisms can be learned to improve model capabilities on specific skills while being composable with out-of-distribution skills, and that independently trained skill neologisms can be composed zero-shot. Finally, we validate zero-shot composition of independently learned skill neologisms on the more realistic natural language setting of the Skill-Mix benchmark. These results suggest that skill neologisms may provide a scalable path towards skill-based continual learning.} }
Endnote
%0 Conference Paper %T Skill Neologisms: Towards Skill-based Continual Learning %A Antonin Berthon %A Nicolás Astorga %A Mihaela Van Der Schaar %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-berthon26a %I PMLR %P 7815--7836 %U https://proceedings.mlr.press/v306/berthon26a.html %V 306 %X Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending model capabilities to new skills in a scalable manner is an open problem: fine-tuning and parameter-efficient variants risk catastrophic forgetting, while context-based approaches have limited expressiveness and are constrained by the model’s effective context. We explore skill neologisms–soft tokens integrated in the model’s vocabulary and optimized to improve capabilities over a specific skill–as a way to selectively acquire new skills without weight updates. We first observe that pretrained LLMs already exhibit tokens associated with procedural knowledge. We then show on a controlled synthetic task that skill neologisms can be learned to improve model capabilities on specific skills while being composable with out-of-distribution skills, and that independently trained skill neologisms can be composed zero-shot. Finally, we validate zero-shot composition of independently learned skill neologisms on the more realistic natural language setting of the Skill-Mix benchmark. These results suggest that skill neologisms may provide a scalable path towards skill-based continual learning.
APA
Berthon, A., Astorga, N. & Van Der Schaar, M.. (2026). Skill Neologisms: Towards Skill-based Continual Learning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:7815-7836 Available from https://proceedings.mlr.press/v306/berthon26a.html.

Related Material