LAMP: Extracting Local Decision Surfaces From Large Language Models

Ryan Chen, Youngmin Ko, Catherine Cho, Zeyu Zhang, Mauro Giuffrè, Sunny Chung, Dennis Shung, Bradly C. Stadie
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4546-4554, 2026.

Abstract

We introduce \textbf{LAMP} (\textbf{L}ocal \textbf{A}ttribution \textbf{M}apping \textbf{P}robe), a method that shines light onto a black-box language model’s decision surface and studies how reliably a model maps its stated reasons to its reported predictions by approximating a decision surface. LAMP treats the model’s own self-reported explanations as a coordinate system and fits a locally linear surrogate that links those weights to the model’s output. By doing so, it reveals how much the stated factors steer the model’s decisions. We apply LAMP to three tasks: \emph{sentiment analysis}, \emph{controversial-topic detection}, and \emph{safety-prompt auditing}. Across these tasks, LAMP reveals that many language models’ locally approximated linear decision landscapes overall agree with human judgments on explanation quality and, on a clinical case-file data set, align with expert assessments. Since LAMP operates without requiring access to model gradients, logits, or internal activations, it serves as a practical and lightweight framework for auditing proprietary language models, and enabling assessment of whether a model appears to behave consistently with the explanations it provides.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-chen26h, title = { LAMP: Extracting Local Decision Surfaces From Large Language Models }, author = {Chen, Ryan and Ko, Youngmin and Cho, Catherine and Zhang, Zeyu and Giuffr{\`e}, Mauro and Chung, Sunny and Shung, Dennis and Stadie, Bradly C.}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4546--4554}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/chen26h/chen26h.pdf}, url = {https://proceedings.mlr.press/v300/chen26h.html}, abstract = { We introduce \textbf{LAMP} (\textbf{L}ocal \textbf{A}ttribution \textbf{M}apping \textbf{P}robe), a method that shines light onto a black-box language model’s decision surface and studies how reliably a model maps its stated reasons to its reported predictions by approximating a decision surface. LAMP treats the model’s own self-reported explanations as a coordinate system and fits a locally linear surrogate that links those weights to the model’s output. By doing so, it reveals how much the stated factors steer the model’s decisions. We apply LAMP to three tasks: \emph{sentiment analysis}, \emph{controversial-topic detection}, and \emph{safety-prompt auditing}. Across these tasks, LAMP reveals that many language models’ locally approximated linear decision landscapes overall agree with human judgments on explanation quality and, on a clinical case-file data set, align with expert assessments. Since LAMP operates without requiring access to model gradients, logits, or internal activations, it serves as a practical and lightweight framework for auditing proprietary language models, and enabling assessment of whether a model appears to behave consistently with the explanations it provides. } }
Endnote
%0 Conference Paper %T LAMP: Extracting Local Decision Surfaces From Large Language Models %A Ryan Chen %A Youngmin Ko %A Catherine Cho %A Zeyu Zhang %A Mauro Giuffrè %A Sunny Chung %A Dennis Shung %A Bradly C. Stadie %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-chen26h %I PMLR %P 4546--4554 %U https://proceedings.mlr.press/v300/chen26h.html %V 300 %X We introduce \textbf{LAMP} (\textbf{L}ocal \textbf{A}ttribution \textbf{M}apping \textbf{P}robe), a method that shines light onto a black-box language model’s decision surface and studies how reliably a model maps its stated reasons to its reported predictions by approximating a decision surface. LAMP treats the model’s own self-reported explanations as a coordinate system and fits a locally linear surrogate that links those weights to the model’s output. By doing so, it reveals how much the stated factors steer the model’s decisions. We apply LAMP to three tasks: \emph{sentiment analysis}, \emph{controversial-topic detection}, and \emph{safety-prompt auditing}. Across these tasks, LAMP reveals that many language models’ locally approximated linear decision landscapes overall agree with human judgments on explanation quality and, on a clinical case-file data set, align with expert assessments. Since LAMP operates without requiring access to model gradients, logits, or internal activations, it serves as a practical and lightweight framework for auditing proprietary language models, and enabling assessment of whether a model appears to behave consistently with the explanations it provides.
APA
Chen, R., Ko, Y., Cho, C., Zhang, Z., Giuffrè, M., Chung, S., Shung, D. & Stadie, B.C.. (2026). LAMP: Extracting Local Decision Surfaces From Large Language Models . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4546-4554 Available from https://proceedings.mlr.press/v300/chen26h.html.

Related Material