Who to Trust? Aggregating Client Predictions in Federated Distillation.

Viktor Kovalchuk, Denis Son, Arman Bolatov, Mohsen Guizani, Samuel Horváth, Maxim Panov, Martin Takáč, Eduard Gorbunov, Nikita Kotelevskii
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3162-3181, 2026.

Abstract

Federated Distillation enables distributed learning for clients with heterogeneous model architectures. In this paradigm, the server and clients exchange predictions on a shared unlabeled public dataset, rather than model parameters or gradients. % However, under data heterogeneity (e.g., class mismatch), clients produce unreliable predictions for instances from unfamiliar classes. % An equally weighted combination of such predictions corrupts the teacher signal used for distillation. % In this paper, we theoretically analyze Federated Distillation and show that aggregating client predictions on a shared public dataset converges to a neighborhood of the optimum, with the neighborhood size controlled by the aggregation quality. % We propose two uncertainty-aware aggregation methods, $\textbf{UWA}$ and $\textbf{sUWA}$, that use density-based estimates to down-weight unreliable client predictions. % Experiments on image and text classification datasets confirm that our methods are most effective under high data heterogeneity, while matching standard averaging when heterogeneity is low.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-kovalchuk26a, title = {Who to Trust? {Aggregating} Client Predictions in Federated Distillation.}, author = {Kovalchuk, Viktor and Son, Denis and Bolatov, Arman and Guizani, Mohsen and Horv\'{a}th, Samuel and Panov, Maxim and Tak\'{a}{\v{c}}, Martin and Gorbunov, Eduard and Kotelevskii, Nikita}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {3162--3181}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/kovalchuk26a/kovalchuk26a.pdf}, url = {https://proceedings.mlr.press/v337/kovalchuk26a.html}, abstract = {Federated Distillation enables distributed learning for clients with heterogeneous model architectures. In this paradigm, the server and clients exchange predictions on a shared unlabeled public dataset, rather than model parameters or gradients. % However, under data heterogeneity (e.g., class mismatch), clients produce unreliable predictions for instances from unfamiliar classes. % An equally weighted combination of such predictions corrupts the teacher signal used for distillation. % In this paper, we theoretically analyze Federated Distillation and show that aggregating client predictions on a shared public dataset converges to a neighborhood of the optimum, with the neighborhood size controlled by the aggregation quality. % We propose two uncertainty-aware aggregation methods, $\textbf{UWA}$ and $\textbf{sUWA}$, that use density-based estimates to down-weight unreliable client predictions. % Experiments on image and text classification datasets confirm that our methods are most effective under high data heterogeneity, while matching standard averaging when heterogeneity is low.} }
Endnote
%0 Conference Paper %T Who to Trust? Aggregating Client Predictions in Federated Distillation. %A Viktor Kovalchuk %A Denis Son %A Arman Bolatov %A Mohsen Guizani %A Samuel Horváth %A Maxim Panov %A Martin Takáč %A Eduard Gorbunov %A Nikita Kotelevskii %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-kovalchuk26a %I PMLR %P 3162--3181 %U https://proceedings.mlr.press/v337/kovalchuk26a.html %V 337 %X Federated Distillation enables distributed learning for clients with heterogeneous model architectures. In this paradigm, the server and clients exchange predictions on a shared unlabeled public dataset, rather than model parameters or gradients. % However, under data heterogeneity (e.g., class mismatch), clients produce unreliable predictions for instances from unfamiliar classes. % An equally weighted combination of such predictions corrupts the teacher signal used for distillation. % In this paper, we theoretically analyze Federated Distillation and show that aggregating client predictions on a shared public dataset converges to a neighborhood of the optimum, with the neighborhood size controlled by the aggregation quality. % We propose two uncertainty-aware aggregation methods, $\textbf{UWA}$ and $\textbf{sUWA}$, that use density-based estimates to down-weight unreliable client predictions. % Experiments on image and text classification datasets confirm that our methods are most effective under high data heterogeneity, while matching standard averaging when heterogeneity is low.
APA
Kovalchuk, V., Son, D., Bolatov, A., Guizani, M., Horváth, S., Panov, M., Takáč, M., Gorbunov, E. & Kotelevskii, N.. (2026). Who to Trust? Aggregating Client Predictions in Federated Distillation.. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:3162-3181 Available from https://proceedings.mlr.press/v337/kovalchuk26a.html.

Related Material