Barriers to Counterfactual Credit Attribution for Autoregressive Models

Aloni Cohen, Chenhao Zhang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:21135-21153, 2026.

Abstract

Generative AI disrupts the practice of giving credit to work that came before. Ideally, a generative model would give credit to any work on which its output depends in a significant way. Counterfactual credit attribution (CCA) is a technical condition formalizing this goal—a relaxation of differential privacy—recently introduced by Livni, Moran, Nissim, and Pabbaraju (2024) who studied it in the PAC learning setting. We initiate the study of CCA generative models. Specifically, we consider autoregressive models giving credit to a deployment-time dataset (e.g., a RAG database). We uncover barriers to two natural approaches to CCA autoregressive models. First, we show that imposing CCA on the underlying next-token predictor does not guarantee that the model is CCA: CCA does not compose autoregressively (unlike DP). Second, we consider a different approach to building CCA models which we call retrofitting. Retrofitting takes a model that does not attribute credit, and adds credit onto it. Given black-box access to the starting model, retrofitting requires query complexity exponential in the length of the model’s outputs.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cohen26e, title = {Barriers to Counterfactual Credit Attribution for Autoregressive Models}, author = {Cohen, Aloni and Zhang, Chenhao}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {21135--21153}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cohen26e/cohen26e.pdf}, url = {https://proceedings.mlr.press/v306/cohen26e.html}, abstract = {Generative AI disrupts the practice of giving credit to work that came before. Ideally, a generative model would give credit to any work on which its output depends in a significant way. Counterfactual credit attribution (CCA) is a technical condition formalizing this goal—a relaxation of differential privacy—recently introduced by Livni, Moran, Nissim, and Pabbaraju (2024) who studied it in the PAC learning setting. We initiate the study of CCA generative models. Specifically, we consider autoregressive models giving credit to a deployment-time dataset (e.g., a RAG database). We uncover barriers to two natural approaches to CCA autoregressive models. First, we show that imposing CCA on the underlying next-token predictor does not guarantee that the model is CCA: CCA does not compose autoregressively (unlike DP). Second, we consider a different approach to building CCA models which we call retrofitting. Retrofitting takes a model that does not attribute credit, and adds credit onto it. Given black-box access to the starting model, retrofitting requires query complexity exponential in the length of the model’s outputs.} }
Endnote
%0 Conference Paper %T Barriers to Counterfactual Credit Attribution for Autoregressive Models %A Aloni Cohen %A Chenhao Zhang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cohen26e %I PMLR %P 21135--21153 %U https://proceedings.mlr.press/v306/cohen26e.html %V 306 %X Generative AI disrupts the practice of giving credit to work that came before. Ideally, a generative model would give credit to any work on which its output depends in a significant way. Counterfactual credit attribution (CCA) is a technical condition formalizing this goal—a relaxation of differential privacy—recently introduced by Livni, Moran, Nissim, and Pabbaraju (2024) who studied it in the PAC learning setting. We initiate the study of CCA generative models. Specifically, we consider autoregressive models giving credit to a deployment-time dataset (e.g., a RAG database). We uncover barriers to two natural approaches to CCA autoregressive models. First, we show that imposing CCA on the underlying next-token predictor does not guarantee that the model is CCA: CCA does not compose autoregressively (unlike DP). Second, we consider a different approach to building CCA models which we call retrofitting. Retrofitting takes a model that does not attribute credit, and adds credit onto it. Given black-box access to the starting model, retrofitting requires query complexity exponential in the length of the model’s outputs.
APA
Cohen, A. & Zhang, C.. (2026). Barriers to Counterfactual Credit Attribution for Autoregressive Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:21135-21153 Available from https://proceedings.mlr.press/v306/cohen26e.html.

Related Material