Explanation Design in Strategic Learning: Sufficient Explanations That Induce Non-harmful Responses

Kiet Q. H. Vo, Siu Lun Chau, Masahiro Kato, Yixin Wang, Krikamol Muandet
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:613-621, 2026.

Abstract

We study the design of explanations in algorithmic decision-making with strategic agents—individuals who may modify their inputs in response to explanations of a decision maker’s (DM’s) predictive model. While the demand for algorithmic transparency has led much prior work to assume full model disclosure, in practice DMs typically provide only partial information via explanations, which can cause agents to misinterpret the model and take actions that unintentionally reduce their own utility. A central open question is therefore how DMs should communicate explanations that avoid harming strategic agents while still supporting their own goals, e.g., minimising predictive error. In this work, we analyse widely used explanation methods and establish a necessary condition to prevent explanations from inducing self-harming responses. Furthermore, we show that action recommendation-based explanations (ARexes), which encompass counterfactual explanations, are sufficient to induce all non-harmful responses. Under a conditional homogeneity assumption, this sufficiency extends to ARex-generating methods, echoing the revelation principle in information design. To demonstrate their practical utility, we introduce a simple learning procedure that jointly optimises the predictive model and the explanation-generating policy. Experiments on both synthetic and real-world tasks show that ARexes enable DMs to achieve high predictive performance while preserving agents’ utility, offering a principled strategy for safe and effective partial model disclosure.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-vo26a, title = { Explanation Design in Strategic Learning: Sufficient Explanations That Induce Non-harmful Responses }, author = {Vo, Kiet Q. H. and Chau, Siu Lun and Kato, Masahiro and Wang, Yixin and Muandet, Krikamol}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {613--621}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/vo26a/vo26a.pdf}, url = {https://proceedings.mlr.press/v300/vo26a.html}, abstract = { We study the design of explanations in algorithmic decision-making with strategic agents—individuals who may modify their inputs in response to explanations of a decision maker’s (DM’s) predictive model. While the demand for algorithmic transparency has led much prior work to assume full model disclosure, in practice DMs typically provide only partial information via explanations, which can cause agents to misinterpret the model and take actions that unintentionally reduce their own utility. A central open question is therefore how DMs should communicate explanations that avoid harming strategic agents while still supporting their own goals, e.g., minimising predictive error. In this work, we analyse widely used explanation methods and establish a necessary condition to prevent explanations from inducing self-harming responses. Furthermore, we show that action recommendation-based explanations (ARexes), which encompass counterfactual explanations, are sufficient to induce all non-harmful responses. Under a conditional homogeneity assumption, this sufficiency extends to ARex-generating methods, echoing the revelation principle in information design. To demonstrate their practical utility, we introduce a simple learning procedure that jointly optimises the predictive model and the explanation-generating policy. Experiments on both synthetic and real-world tasks show that ARexes enable DMs to achieve high predictive performance while preserving agents’ utility, offering a principled strategy for safe and effective partial model disclosure. } }
Endnote
%0 Conference Paper %T Explanation Design in Strategic Learning: Sufficient Explanations That Induce Non-harmful Responses %A Kiet Q. H. Vo %A Siu Lun Chau %A Masahiro Kato %A Yixin Wang %A Krikamol Muandet %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-vo26a %I PMLR %P 613--621 %U https://proceedings.mlr.press/v300/vo26a.html %V 300 %X We study the design of explanations in algorithmic decision-making with strategic agents—individuals who may modify their inputs in response to explanations of a decision maker’s (DM’s) predictive model. While the demand for algorithmic transparency has led much prior work to assume full model disclosure, in practice DMs typically provide only partial information via explanations, which can cause agents to misinterpret the model and take actions that unintentionally reduce their own utility. A central open question is therefore how DMs should communicate explanations that avoid harming strategic agents while still supporting their own goals, e.g., minimising predictive error. In this work, we analyse widely used explanation methods and establish a necessary condition to prevent explanations from inducing self-harming responses. Furthermore, we show that action recommendation-based explanations (ARexes), which encompass counterfactual explanations, are sufficient to induce all non-harmful responses. Under a conditional homogeneity assumption, this sufficiency extends to ARex-generating methods, echoing the revelation principle in information design. To demonstrate their practical utility, we introduce a simple learning procedure that jointly optimises the predictive model and the explanation-generating policy. Experiments on both synthetic and real-world tasks show that ARexes enable DMs to achieve high predictive performance while preserving agents’ utility, offering a principled strategy for safe and effective partial model disclosure.
APA
Vo, K.Q.H., Chau, S.L., Kato, M., Wang, Y. & Muandet, K.. (2026). Explanation Design in Strategic Learning: Sufficient Explanations That Induce Non-harmful Responses . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:613-621 Available from https://proceedings.mlr.press/v300/vo26a.html.

Related Material