ARGUS: Argumentation-Based Minimal-Change Repair for Verifiable LLM Self-Explanations

Yifan Xiao, Shijie Li
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7536-7551, 2026.

Abstract

When large language models produce natural-language rationales, those explanations are frequently unfaithful to the model’s actual reasoning—and no existing framework provides a principled way to repair them when new evidence arrives. We introduce ARGUS, a framework that structures {LLM} self-explanations as Dung-style abstract argumentation frameworks and verifies them under grounded and preferred semantics. When new or uncertain evidence renders an explanation inconsistent, ARGUS computes a minimum-cost set of edit operations that restores the desired acceptability status of the target argument. The repair operator satisfies adapted AGM revision postulates and is bidirectionally characterized by them (Representation Theorem): the decision problem is in P under grounded semantics, NP-complete under preferred and stable semantics, and $\Sigma_2^P$-complete under skeptical stable semantics. A $k$-neighborhood approximation and an answer set programming (ASP) encoding ensure scalability to practical framework sizes. We validate the framework on HotpotQA and FEVER, where ARGUS achieves relative improvements of 10.3% in faithfulness and 14.5% in contestability over the strongest argumentation baseline while requiring fewer repair operations than all repair-capable competing methods.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-xiao26a, title = {ARGUS: Argumentation-Based Minimal-Change Repair for Verifiable {LLM} Self-Explanations}, author = {Xiao, Yifan and Li, Shijie}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {7536--7551}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/xiao26a/xiao26a.pdf}, url = {https://proceedings.mlr.press/v337/xiao26a.html}, abstract = {When large language models produce natural-language rationales, those explanations are frequently unfaithful to the model’s actual reasoning—and no existing framework provides a principled way to repair them when new evidence arrives. We introduce ARGUS, a framework that structures {LLM} self-explanations as Dung-style abstract argumentation frameworks and verifies them under grounded and preferred semantics. When new or uncertain evidence renders an explanation inconsistent, ARGUS computes a minimum-cost set of edit operations that restores the desired acceptability status of the target argument. The repair operator satisfies adapted AGM revision postulates and is bidirectionally characterized by them (Representation Theorem): the decision problem is in P under grounded semantics, NP-complete under preferred and stable semantics, and $\Sigma_2^P$-complete under skeptical stable semantics. A $k$-neighborhood approximation and an answer set programming (ASP) encoding ensure scalability to practical framework sizes. We validate the framework on HotpotQA and FEVER, where ARGUS achieves relative improvements of 10.3% in faithfulness and 14.5% in contestability over the strongest argumentation baseline while requiring fewer repair operations than all repair-capable competing methods.} }
Endnote
%0 Conference Paper %T ARGUS: Argumentation-Based Minimal-Change Repair for Verifiable LLM Self-Explanations %A Yifan Xiao %A Shijie Li %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-xiao26a %I PMLR %P 7536--7551 %U https://proceedings.mlr.press/v337/xiao26a.html %V 337 %X When large language models produce natural-language rationales, those explanations are frequently unfaithful to the model’s actual reasoning—and no existing framework provides a principled way to repair them when new evidence arrives. We introduce ARGUS, a framework that structures {LLM} self-explanations as Dung-style abstract argumentation frameworks and verifies them under grounded and preferred semantics. When new or uncertain evidence renders an explanation inconsistent, ARGUS computes a minimum-cost set of edit operations that restores the desired acceptability status of the target argument. The repair operator satisfies adapted AGM revision postulates and is bidirectionally characterized by them (Representation Theorem): the decision problem is in P under grounded semantics, NP-complete under preferred and stable semantics, and $\Sigma_2^P$-complete under skeptical stable semantics. A $k$-neighborhood approximation and an answer set programming (ASP) encoding ensure scalability to practical framework sizes. We validate the framework on HotpotQA and FEVER, where ARGUS achieves relative improvements of 10.3% in faithfulness and 14.5% in contestability over the strongest argumentation baseline while requiring fewer repair operations than all repair-capable competing methods.
APA
Xiao, Y. & Li, S.. (2026). ARGUS: Argumentation-Based Minimal-Change Repair for Verifiable LLM Self-Explanations. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:7536-7551 Available from https://proceedings.mlr.press/v337/xiao26a.html.

Related Material