Disentangling Federated Learning Heterogeneity: A Dual-Perspective Analysis of Quantifying Skew Versus Scarcity

Wenkai Zeng, NAN YANG, Zhiyu Zhu, Zhibo Jin, Dong Yuan
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:5068-5076, 2026.

Abstract

Federated Learning faces significant challenges due to data heterogeneity, which manifests as Label Distribution Skew and label missingness. We propose Skew-Scarcity Disentanglement Indicator (SSDI), a novel metric that decomposes heterogeneity into two disentangled components: Label Distribution Skew (LDS) (quantity skew of present labels) and Label Coverage Deficiency (LCD) (deviation due to missing labels). Using a PAC-Bayesian framework, we derive a generalization bound indicating that Label Coverage Deficiency becomes the dominant risk factor as the number of clients increases, severely degrading accuracy on rare labels. Our study reveals that, for a fixed number of labels, increasing clients is a primary driver of per-label accuracy variance by exacerbating Label Coverage Deficiency. Moreover, a higher global missing rate intensifies this divergence effect and can precipitate severe performance breakdown at a lower critical threshold of clients. Experiments on vision benchmarks confirm that SSDI accurately captures the severity of performance divergence. The SSDI framework provides a principled tool for diagnosing heterogeneity and guiding targeted mitigation strategies. The code for the SSDI-controlled client-label matrix generation used in our experiments is available at \url{https://github.com/wkzeng/SSDI.git.}

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-zeng26b, title = { Disentangling Federated Learning Heterogeneity: A Dual-Perspective Analysis of Quantifying Skew Versus Scarcity }, author = {Zeng, Wenkai and YANG, NAN and Zhu, Zhiyu and Jin, Zhibo and Yuan, Dong}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {5068--5076}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/zeng26b/zeng26b.pdf}, url = {https://proceedings.mlr.press/v300/zeng26b.html}, abstract = { Federated Learning faces significant challenges due to data heterogeneity, which manifests as Label Distribution Skew and label missingness. We propose Skew-Scarcity Disentanglement Indicator (SSDI), a novel metric that decomposes heterogeneity into two disentangled components: Label Distribution Skew (LDS) (quantity skew of present labels) and Label Coverage Deficiency (LCD) (deviation due to missing labels). Using a PAC-Bayesian framework, we derive a generalization bound indicating that Label Coverage Deficiency becomes the dominant risk factor as the number of clients increases, severely degrading accuracy on rare labels. Our study reveals that, for a fixed number of labels, increasing clients is a primary driver of per-label accuracy variance by exacerbating Label Coverage Deficiency. Moreover, a higher global missing rate intensifies this divergence effect and can precipitate severe performance breakdown at a lower critical threshold of clients. Experiments on vision benchmarks confirm that SSDI accurately captures the severity of performance divergence. The SSDI framework provides a principled tool for diagnosing heterogeneity and guiding targeted mitigation strategies. The code for the SSDI-controlled client-label matrix generation used in our experiments is available at \url{https://github.com/wkzeng/SSDI.git.} } }
Endnote
%0 Conference Paper %T Disentangling Federated Learning Heterogeneity: A Dual-Perspective Analysis of Quantifying Skew Versus Scarcity %A Wenkai Zeng %A NAN YANG %A Zhiyu Zhu %A Zhibo Jin %A Dong Yuan %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-zeng26b %I PMLR %P 5068--5076 %U https://proceedings.mlr.press/v300/zeng26b.html %V 300 %X Federated Learning faces significant challenges due to data heterogeneity, which manifests as Label Distribution Skew and label missingness. We propose Skew-Scarcity Disentanglement Indicator (SSDI), a novel metric that decomposes heterogeneity into two disentangled components: Label Distribution Skew (LDS) (quantity skew of present labels) and Label Coverage Deficiency (LCD) (deviation due to missing labels). Using a PAC-Bayesian framework, we derive a generalization bound indicating that Label Coverage Deficiency becomes the dominant risk factor as the number of clients increases, severely degrading accuracy on rare labels. Our study reveals that, for a fixed number of labels, increasing clients is a primary driver of per-label accuracy variance by exacerbating Label Coverage Deficiency. Moreover, a higher global missing rate intensifies this divergence effect and can precipitate severe performance breakdown at a lower critical threshold of clients. Experiments on vision benchmarks confirm that SSDI accurately captures the severity of performance divergence. The SSDI framework provides a principled tool for diagnosing heterogeneity and guiding targeted mitigation strategies. The code for the SSDI-controlled client-label matrix generation used in our experiments is available at \url{https://github.com/wkzeng/SSDI.git.}
APA
Zeng, W., YANG, N., Zhu, Z., Jin, Z. & Yuan, D.. (2026). Disentangling Federated Learning Heterogeneity: A Dual-Perspective Analysis of Quantifying Skew Versus Scarcity . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:5068-5076 Available from https://proceedings.mlr.press/v300/zeng26b.html.

Related Material