Evaluating Contextual Illegality: AI Compliance in Corporate Law Scenarios

Hilal Aka, Joe Kwon, Noam Kolt
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:1499-1527, 2026.

Abstract

While AI models often refuse explicitly unlawful requests, in real-world scenarios illegality often depends on context. We evaluate frontier models on contextual illegality across four corporate law scenarios in which routine actions—editing documents, trading stock, requesting payment, approving communications—become unlawful due to circumstances such as pending investigations or bankruptcy filings. We study both chat and agentic settings and compare results to a human baseline. The best-performing models consistently followed lawful requests and refused unlawful requests, though performance varied substantially between different scenarios and models. We also identify distinct failure modes, such as excessive refusal of lawful requests, and find higher performance in reasoning models and agentic environments. By studying contextual illegality in these controlled environments, we develop a methodology that can be extended to evaluate the legal compliance of AI models in additional scenarios and domains.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-aka26a, title = {Evaluating Contextual Illegality: {AI} Compliance in Corporate Law Scenarios}, author = {Aka, Hilal and Kwon, Joe and Kolt, Noam}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {1499--1527}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/aka26a/aka26a.pdf}, url = {https://proceedings.mlr.press/v306/aka26a.html}, abstract = {While AI models often refuse explicitly unlawful requests, in real-world scenarios illegality often depends on context. We evaluate frontier models on contextual illegality across four corporate law scenarios in which routine actions—editing documents, trading stock, requesting payment, approving communications—become unlawful due to circumstances such as pending investigations or bankruptcy filings. We study both chat and agentic settings and compare results to a human baseline. The best-performing models consistently followed lawful requests and refused unlawful requests, though performance varied substantially between different scenarios and models. We also identify distinct failure modes, such as excessive refusal of lawful requests, and find higher performance in reasoning models and agentic environments. By studying contextual illegality in these controlled environments, we develop a methodology that can be extended to evaluate the legal compliance of AI models in additional scenarios and domains.} }
Endnote
%0 Conference Paper %T Evaluating Contextual Illegality: AI Compliance in Corporate Law Scenarios %A Hilal Aka %A Joe Kwon %A Noam Kolt %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-aka26a %I PMLR %P 1499--1527 %U https://proceedings.mlr.press/v306/aka26a.html %V 306 %X While AI models often refuse explicitly unlawful requests, in real-world scenarios illegality often depends on context. We evaluate frontier models on contextual illegality across four corporate law scenarios in which routine actions—editing documents, trading stock, requesting payment, approving communications—become unlawful due to circumstances such as pending investigations or bankruptcy filings. We study both chat and agentic settings and compare results to a human baseline. The best-performing models consistently followed lawful requests and refused unlawful requests, though performance varied substantially between different scenarios and models. We also identify distinct failure modes, such as excessive refusal of lawful requests, and find higher performance in reasoning models and agentic environments. By studying contextual illegality in these controlled environments, we develop a methodology that can be extended to evaluate the legal compliance of AI models in additional scenarios and domains.
APA
Aka, H., Kwon, J. & Kolt, N.. (2026). Evaluating Contextual Illegality: AI Compliance in Corporate Law Scenarios. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:1499-1527 Available from https://proceedings.mlr.press/v306/aka26a.html.

Related Material