FIELDING: Clustered Federated Learning with Data Drift

Minghao Li, Dmitrii Avdiukhin, Rana Shahout, Nikita Ivkin, Vladimir Braverman, Minlan Yu
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:811-819, 2026.

Abstract

Federated Learning (FL) trains deep models across edge devices without centralizing raw data. However, client heterogeneity slows down convergence and limits global model accuracy. Clustered FL (CFL) mitigates this by grouping clients with similar representations and training a separate model for each cluster. In practice, client data evolves over time – a phenomenon we refer to as data drift – which breaks cluster homogeneity and degrades performance. Data drift can take different forms depending on whether changes occur in the output values, the input features, or the relationship between them. We propose FIELDING, a CFL framework for handling diverse types of data drift with low overhead. FIELDING detects drift at individual clients and performs selective re-clustering to balance cluster quality and model performance, while remaining robust to varying levels of heterogeneity. Experiments show that FIELDING improves final model accuracy by 2.4–6.9% and achieves target accuracy 1.38x–3.10x faster than existing state-of-the-art CFL methods.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-li26c, title = { FIELDING: Clustered Federated Learning with Data Drift }, author = {Li, Minghao and Avdiukhin, Dmitrii and Shahout, Rana and Ivkin, Nikita and Braverman, Vladimir and Yu, Minlan}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {811--819}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/li26c/li26c.pdf}, url = {https://proceedings.mlr.press/v300/li26c.html}, abstract = { Federated Learning (FL) trains deep models across edge devices without centralizing raw data. However, client heterogeneity slows down convergence and limits global model accuracy. Clustered FL (CFL) mitigates this by grouping clients with similar representations and training a separate model for each cluster. In practice, client data evolves over time – a phenomenon we refer to as data drift – which breaks cluster homogeneity and degrades performance. Data drift can take different forms depending on whether changes occur in the output values, the input features, or the relationship between them. We propose FIELDING, a CFL framework for handling diverse types of data drift with low overhead. FIELDING detects drift at individual clients and performs selective re-clustering to balance cluster quality and model performance, while remaining robust to varying levels of heterogeneity. Experiments show that FIELDING improves final model accuracy by 2.4–6.9% and achieves target accuracy 1.38x–3.10x faster than existing state-of-the-art CFL methods. } }
Endnote
%0 Conference Paper %T FIELDING: Clustered Federated Learning with Data Drift %A Minghao Li %A Dmitrii Avdiukhin %A Rana Shahout %A Nikita Ivkin %A Vladimir Braverman %A Minlan Yu %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-li26c %I PMLR %P 811--819 %U https://proceedings.mlr.press/v300/li26c.html %V 300 %X Federated Learning (FL) trains deep models across edge devices without centralizing raw data. However, client heterogeneity slows down convergence and limits global model accuracy. Clustered FL (CFL) mitigates this by grouping clients with similar representations and training a separate model for each cluster. In practice, client data evolves over time – a phenomenon we refer to as data drift – which breaks cluster homogeneity and degrades performance. Data drift can take different forms depending on whether changes occur in the output values, the input features, or the relationship between them. We propose FIELDING, a CFL framework for handling diverse types of data drift with low overhead. FIELDING detects drift at individual clients and performs selective re-clustering to balance cluster quality and model performance, while remaining robust to varying levels of heterogeneity. Experiments show that FIELDING improves final model accuracy by 2.4–6.9% and achieves target accuracy 1.38x–3.10x faster than existing state-of-the-art CFL methods.
APA
Li, M., Avdiukhin, D., Shahout, R., Ivkin, N., Braverman, V. & Yu, M.. (2026). FIELDING: Clustered Federated Learning with Data Drift . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:811-819 Available from https://proceedings.mlr.press/v300/li26c.html.

Related Material