Capacity and Redundancy Trade-offs in Multi-Task Learning

Asif Khan
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2934-2957, 2026.

Abstract

In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence of limited shared capacity and weak task redundancy. We investigate this effect through a Capacity–Redundancy ({CR}) identity that decomposes the sum of per-task predictive informations into joint predictive information that includes label redundancy defined via total correlation, and a residual coupling term that quantifies interference left unresolved by the shared representation. Additionally, we show two key results: (i) a clustering-gap decomposition that gives a necessary and sufficient condition for clustered sharing to outperform global sharing, and (ii) a gradient–TC bridge in a {Gaussian} multi-task model that formally justifies gradient cosine similarity as a proxy for redundancy ordering. Empirically, we estimate $\Delta$ from validation residual correlations, showing that clustered {LoRA} substantially reduces $\widehat{\Delta}$, outperforms size-matched random partitions, and results in statistically significant gains with multi-seed confidence intervals.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-khan26a, title = {Capacity and Redundancy Trade-offs in Multi-Task Learning}, author = {Khan, Asif}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {2934--2957}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/khan26a/khan26a.pdf}, url = {https://proceedings.mlr.press/v337/khan26a.html}, abstract = {In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence of limited shared capacity and weak task redundancy. We investigate this effect through a Capacity–Redundancy ({CR}) identity that decomposes the sum of per-task predictive informations into joint predictive information that includes label redundancy defined via total correlation, and a residual coupling term that quantifies interference left unresolved by the shared representation. Additionally, we show two key results: (i) a clustering-gap decomposition that gives a necessary and sufficient condition for clustered sharing to outperform global sharing, and (ii) a gradient–TC bridge in a {Gaussian} multi-task model that formally justifies gradient cosine similarity as a proxy for redundancy ordering. Empirically, we estimate $\Delta$ from validation residual correlations, showing that clustered {LoRA} substantially reduces $\widehat{\Delta}$, outperforms size-matched random partitions, and results in statistically significant gains with multi-seed confidence intervals.} }
Endnote
%0 Conference Paper %T Capacity and Redundancy Trade-offs in Multi-Task Learning %A Asif Khan %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-khan26a %I PMLR %P 2934--2957 %U https://proceedings.mlr.press/v337/khan26a.html %V 337 %X In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence of limited shared capacity and weak task redundancy. We investigate this effect through a Capacity–Redundancy ({CR}) identity that decomposes the sum of per-task predictive informations into joint predictive information that includes label redundancy defined via total correlation, and a residual coupling term that quantifies interference left unresolved by the shared representation. Additionally, we show two key results: (i) a clustering-gap decomposition that gives a necessary and sufficient condition for clustered sharing to outperform global sharing, and (ii) a gradient–TC bridge in a {Gaussian} multi-task model that formally justifies gradient cosine similarity as a proxy for redundancy ordering. Empirically, we estimate $\Delta$ from validation residual correlations, showing that clustered {LoRA} substantially reduces $\widehat{\Delta}$, outperforms size-matched random partitions, and results in statistically significant gains with multi-seed confidence intervals.
APA
Khan, A.. (2026). Capacity and Redundancy Trade-offs in Multi-Task Learning. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:2934-2957 Available from https://proceedings.mlr.press/v337/khan26a.html.

Related Material