Corruption-robust Offline Multi-agent Reinforcement Learning from Human Feedback

Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban, Adish Singla, Goran Radanovic
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:1864-1872, 2026.

Abstract

We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset $D$ of trajectory–preference tuples (each preference being an $n$-dimensional binary label vector representing each of the $n$ agents’ preferences), an $\epsilon$-fraction of the samples may be arbitrarily corrupted. We model the problem using the framework of linear Markov games. First, under a \emph{uniform coverage} assumption—where every policy of interest is sufficiently represented in the clean (prior to corruption) data—we introduce a robust estimator that guarantees an $O(\epsilon^{1-o(1)})$ bound on the Nash-equilibrium gap. Next, we move to the more challenging \emph{unilateral coverage} setting, in which only a Nash equilibrium and its single-player deviations are covered: here our proposed algorithm achieves an $O(\sqrt{\epsilon})$ Nash-gap bound. Both of these procedures, however, suffer from intractable computation. To address this, we relax our solution concept to \emph{coarse correlated equilibria} (CCE). Under the same unilateral-coverage regime, we then derive a quasi-polynomial-time algorithm whose CCE gap scales as $O(\sqrt{\epsilon})$. To the best of our knowledge, this is the first systematic treatment of adversarial data corruption in offline MARLHF.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-nika26a, title = { Corruption-robust Offline Multi-agent Reinforcement Learning from Human Feedback }, author = {Nika, Andi and Mandal, Debmalya and Kamalaruban, Parameswaran and Singla, Adish and Radanovic, Goran}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {1864--1872}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/nika26a/nika26a.pdf}, url = {https://proceedings.mlr.press/v300/nika26a.html}, abstract = { We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset $D$ of trajectory–preference tuples (each preference being an $n$-dimensional binary label vector representing each of the $n$ agents’ preferences), an $\epsilon$-fraction of the samples may be arbitrarily corrupted. We model the problem using the framework of linear Markov games. First, under a \emph{uniform coverage} assumption—where every policy of interest is sufficiently represented in the clean (prior to corruption) data—we introduce a robust estimator that guarantees an $O(\epsilon^{1-o(1)})$ bound on the Nash-equilibrium gap. Next, we move to the more challenging \emph{unilateral coverage} setting, in which only a Nash equilibrium and its single-player deviations are covered: here our proposed algorithm achieves an $O(\sqrt{\epsilon})$ Nash-gap bound. Both of these procedures, however, suffer from intractable computation. To address this, we relax our solution concept to \emph{coarse correlated equilibria} (CCE). Under the same unilateral-coverage regime, we then derive a quasi-polynomial-time algorithm whose CCE gap scales as $O(\sqrt{\epsilon})$. To the best of our knowledge, this is the first systematic treatment of adversarial data corruption in offline MARLHF. } }
Endnote
%0 Conference Paper %T Corruption-robust Offline Multi-agent Reinforcement Learning from Human Feedback %A Andi Nika %A Debmalya Mandal %A Parameswaran Kamalaruban %A Adish Singla %A Goran Radanovic %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-nika26a %I PMLR %P 1864--1872 %U https://proceedings.mlr.press/v300/nika26a.html %V 300 %X We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset $D$ of trajectory–preference tuples (each preference being an $n$-dimensional binary label vector representing each of the $n$ agents’ preferences), an $\epsilon$-fraction of the samples may be arbitrarily corrupted. We model the problem using the framework of linear Markov games. First, under a \emph{uniform coverage} assumption—where every policy of interest is sufficiently represented in the clean (prior to corruption) data—we introduce a robust estimator that guarantees an $O(\epsilon^{1-o(1)})$ bound on the Nash-equilibrium gap. Next, we move to the more challenging \emph{unilateral coverage} setting, in which only a Nash equilibrium and its single-player deviations are covered: here our proposed algorithm achieves an $O(\sqrt{\epsilon})$ Nash-gap bound. Both of these procedures, however, suffer from intractable computation. To address this, we relax our solution concept to \emph{coarse correlated equilibria} (CCE). Under the same unilateral-coverage regime, we then derive a quasi-polynomial-time algorithm whose CCE gap scales as $O(\sqrt{\epsilon})$. To the best of our knowledge, this is the first systematic treatment of adversarial data corruption in offline MARLHF.
APA
Nika, A., Mandal, D., Kamalaruban, P., Singla, A. & Radanovic, G.. (2026). Corruption-robust Offline Multi-agent Reinforcement Learning from Human Feedback . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:1864-1872 Available from https://proceedings.mlr.press/v300/nika26a.html.

Related Material