LAVA: A Unified Framework for Finetuning Language and Vision Models

Daorui Ding, Fanhua Shang, Tiancan Feng, Junkang Liu, Hongying Liu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:25130-25151, 2026.

Abstract

LoRA and its variants have attracted considerable attention because of their abilities to tune a negligible number of parameters while achieving comparable downstream performance. This success is largely attributed to the intrinsic low-rank structure of model parameter spaces, which allows LoRA to train two projection matrices to project weights into a low-dimensional subspace and then map them back. However, it does not consider how to explore this low-rank subspace sufficiently and may lose the expression ability accordingly. Moreover, when using LoRA to tune convolution layers, a flatten operation is required to convert tensors into matrices. We argue that this will degrade the model’s performance. In this paper, we address this issue from a general parameter sub-space perspective: we present a unified Language And Vision Adaption finetuning framework (called LAVA). Specifically, we verify the existence of low-rank subspaces in convolution layers empirically and propose to parameterize the increment of both convolution kernels and matrices as sum of learnable rank-1 components. To improve training stability, we analyze the optimization dynamics of LoRA and incorporate orthogonal regularization into our parameterization, for which we give theoretical proof that it will help reduce the variance of the gradient. We conduct various experiments on different downstreaming tasks to validate LAVA’s superiority. For example, when tuning LLaMA2-7b for commonsense tasks, the performance of our LAVA is +1.9% higher than that of LoRA. For metric depth estimation tasks, LAVA only tunes $\sim$1.5% of Depth-Anything (335.3M), and achieves +3.5% $\delta_1$ accuracy against that of LoRA and +5.6% $\delta_1$ accuracy against that of SVDiff.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-ding26o, title = {{LAVA}: A Unified Framework for Finetuning Language and Vision Models}, author = {Ding, Daorui and Shang, Fanhua and Feng, Tiancan and Liu, Junkang and Liu, Hongying}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {25130--25151}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ding26o/ding26o.pdf}, url = {https://proceedings.mlr.press/v306/ding26o.html}, abstract = {LoRA and its variants have attracted considerable attention because of their abilities to tune a negligible number of parameters while achieving comparable downstream performance. This success is largely attributed to the intrinsic low-rank structure of model parameter spaces, which allows LoRA to train two projection matrices to project weights into a low-dimensional subspace and then map them back. However, it does not consider how to explore this low-rank subspace sufficiently and may lose the expression ability accordingly. Moreover, when using LoRA to tune convolution layers, a flatten operation is required to convert tensors into matrices. We argue that this will degrade the model’s performance. In this paper, we address this issue from a general parameter sub-space perspective: we present a unified Language And Vision Adaption finetuning framework (called LAVA). Specifically, we verify the existence of low-rank subspaces in convolution layers empirically and propose to parameterize the increment of both convolution kernels and matrices as sum of learnable rank-1 components. To improve training stability, we analyze the optimization dynamics of LoRA and incorporate orthogonal regularization into our parameterization, for which we give theoretical proof that it will help reduce the variance of the gradient. We conduct various experiments on different downstreaming tasks to validate LAVA’s superiority. For example, when tuning LLaMA2-7b for commonsense tasks, the performance of our LAVA is +1.9% higher than that of LoRA. For metric depth estimation tasks, LAVA only tunes $\sim$1.5% of Depth-Anything (335.3M), and achieves +3.5% $\delta_1$ accuracy against that of LoRA and +5.6% $\delta_1$ accuracy against that of SVDiff.} }
Endnote
%0 Conference Paper %T LAVA: A Unified Framework for Finetuning Language and Vision Models %A Daorui Ding %A Fanhua Shang %A Tiancan Feng %A Junkang Liu %A Hongying Liu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-ding26o %I PMLR %P 25130--25151 %U https://proceedings.mlr.press/v306/ding26o.html %V 306 %X LoRA and its variants have attracted considerable attention because of their abilities to tune a negligible number of parameters while achieving comparable downstream performance. This success is largely attributed to the intrinsic low-rank structure of model parameter spaces, which allows LoRA to train two projection matrices to project weights into a low-dimensional subspace and then map them back. However, it does not consider how to explore this low-rank subspace sufficiently and may lose the expression ability accordingly. Moreover, when using LoRA to tune convolution layers, a flatten operation is required to convert tensors into matrices. We argue that this will degrade the model’s performance. In this paper, we address this issue from a general parameter sub-space perspective: we present a unified Language And Vision Adaption finetuning framework (called LAVA). Specifically, we verify the existence of low-rank subspaces in convolution layers empirically and propose to parameterize the increment of both convolution kernels and matrices as sum of learnable rank-1 components. To improve training stability, we analyze the optimization dynamics of LoRA and incorporate orthogonal regularization into our parameterization, for which we give theoretical proof that it will help reduce the variance of the gradient. We conduct various experiments on different downstreaming tasks to validate LAVA’s superiority. For example, when tuning LLaMA2-7b for commonsense tasks, the performance of our LAVA is +1.9% higher than that of LoRA. For metric depth estimation tasks, LAVA only tunes $\sim$1.5% of Depth-Anything (335.3M), and achieves +3.5% $\delta_1$ accuracy against that of LoRA and +5.6% $\delta_1$ accuracy against that of SVDiff.
APA
Ding, D., Shang, F., Feng, T., Liu, J. & Liu, H.. (2026). LAVA: A Unified Framework for Finetuning Language and Vision Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:25130-25151 Available from https://proceedings.mlr.press/v306/ding26o.html.

Related Material