Weight Updates as Activation Shifts: A Principled Framework for Steering

Dyah Adila, John Cooper, Alexander Yun, Avi Trost, Frederic Sala
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:565-590, 2026.

Abstract

Activation steering promises to be an extremely parameter-efficient form of adaptation, but its effectiveness depends on critical design choices—such as intervention location and parameterization—that currently rely on empirical heuristics rather than a principled foundation. We establish a first-order equivalence between activation-space interventions and weight-space updates, deriving the conditions under which activation steering can replicate fine-tuning behavior. This equivalence yields a principled framework for steering design and identifies the post-block output as a theoretically-backed and highly expressive intervention site. We explain why certain intervention locations outperform others and show that weight updates and activation updates play distinct, complementary functional roles. This analysis motivates a new approach—joint adaptation—that trains in both spaces simultaneously. Our post-block steering achieves accuracy within $0.2%\text{–}0.9%$ of full-parameter tuning, on average across tasks and models, while training only $0.04%$ of model parameters. It consistently outperforms prior activation steering methods such as ReFT and PEFT approaches including LoRA, while using significantly fewer parameters. Finally, we show that joint adaptation often surpasses the performance ceilings of weight and activation updates in isolation, introducing a new paradigm for efficient model adaptation

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-adila26a, title = {Weight Updates as Activation Shifts: A Principled Framework for Steering}, author = {Adila, Dyah and Cooper, John and Yun, Alexander and Trost, Avi and Sala, Frederic}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {565--590}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/adila26a/adila26a.pdf}, url = {https://proceedings.mlr.press/v306/adila26a.html}, abstract = {Activation steering promises to be an extremely parameter-efficient form of adaptation, but its effectiveness depends on critical design choices—such as intervention location and parameterization—that currently rely on empirical heuristics rather than a principled foundation. We establish a first-order equivalence between activation-space interventions and weight-space updates, deriving the conditions under which activation steering can replicate fine-tuning behavior. This equivalence yields a principled framework for steering design and identifies the post-block output as a theoretically-backed and highly expressive intervention site. We explain why certain intervention locations outperform others and show that weight updates and activation updates play distinct, complementary functional roles. This analysis motivates a new approach—joint adaptation—that trains in both spaces simultaneously. Our post-block steering achieves accuracy within $0.2%\text{–}0.9%$ of full-parameter tuning, on average across tasks and models, while training only $0.04%$ of model parameters. It consistently outperforms prior activation steering methods such as ReFT and PEFT approaches including LoRA, while using significantly fewer parameters. Finally, we show that joint adaptation often surpasses the performance ceilings of weight and activation updates in isolation, introducing a new paradigm for efficient model adaptation} }
Endnote
%0 Conference Paper %T Weight Updates as Activation Shifts: A Principled Framework for Steering %A Dyah Adila %A John Cooper %A Alexander Yun %A Avi Trost %A Frederic Sala %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-adila26a %I PMLR %P 565--590 %U https://proceedings.mlr.press/v306/adila26a.html %V 306 %X Activation steering promises to be an extremely parameter-efficient form of adaptation, but its effectiveness depends on critical design choices—such as intervention location and parameterization—that currently rely on empirical heuristics rather than a principled foundation. We establish a first-order equivalence between activation-space interventions and weight-space updates, deriving the conditions under which activation steering can replicate fine-tuning behavior. This equivalence yields a principled framework for steering design and identifies the post-block output as a theoretically-backed and highly expressive intervention site. We explain why certain intervention locations outperform others and show that weight updates and activation updates play distinct, complementary functional roles. This analysis motivates a new approach—joint adaptation—that trains in both spaces simultaneously. Our post-block steering achieves accuracy within $0.2%\text{–}0.9%$ of full-parameter tuning, on average across tasks and models, while training only $0.04%$ of model parameters. It consistently outperforms prior activation steering methods such as ReFT and PEFT approaches including LoRA, while using significantly fewer parameters. Finally, we show that joint adaptation often surpasses the performance ceilings of weight and activation updates in isolation, introducing a new paradigm for efficient model adaptation
APA
Adila, D., Cooper, J., Yun, A., Trost, A. & Sala, F.. (2026). Weight Updates as Activation Shifts: A Principled Framework for Steering. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:565-590 Available from https://proceedings.mlr.press/v306/adila26a.html.

Related Material