On Efficient Scaling of GNNs via IO-Aware Layers Implementations

Daria Fomina, Daniil Krasylnikov, Alexey Boykov, Andrey Dolgovyazov, Vyacheslav Zhdanovskiy, Fedor Velikonivtsev
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:31329-31376, 2026.

Abstract

Graph Neural Networks (GNNs) are bottlenecked by sparse, irregular memory access. Popular frameworks such as DGL and PyTorch Geometric support general message passing, but complex layers often materialize edge-wise intermediates, increasing memory traffic and limiting scalability on large graphs. We take an I/O- and arithmetic-intensity–centric view and show that widely used layers fall into three kernel families: SpMM-based convolutions, reduction-based aggregations, and attention-based layers (GATv2/Graph Transformer). For each family, we develop GPU kernels that reduce data movement, improve locality, and remain robust across realistic graphs. We also study graph reordering and find that its impact depends on the kernel mapping: it benefits neighbor-parallel (gather-dominated) kernels more consistently than feature-parallel designs. Empirically, our fused attention kernels reach up to 3.9$\times$ speedup for Graph Transformer (median 1.6$\times$), with Tensor Core (block-sparse) variants up to 7.3$\times$ on locally dense graphs; for GATv2 we reach up to 8.5$\times$ speedup (median 2.0$\times$) while reducing peak memory by up to 76$\times$ (median 6$\times$). Our degree-aware reduction kernels achieve up to 10$\times$ speedup (median 2.6$\times$). For SpMM-based layers, properly cached cuSPARSE achieves up to 8$\times$ speedup over DGL and outperforms evaluated custom baselines in the majority of evaluations. We release our implementations as drop-in replacements in our GitHub repository to support reproducible, hardware-aware GNN acceleration.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-fomina26a, title = {On Efficient Scaling of {GNN}s via {IO}-Aware Layers Implementations}, author = {Fomina, Daria and Krasylnikov, Daniil and Boykov, Alexey and Dolgovyazov, Andrey and Zhdanovskiy, Vyacheslav and Velikonivtsev, Fedor}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {31329--31376}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/fomina26a/fomina26a.pdf}, url = {https://proceedings.mlr.press/v306/fomina26a.html}, abstract = {Graph Neural Networks (GNNs) are bottlenecked by sparse, irregular memory access. Popular frameworks such as DGL and PyTorch Geometric support general message passing, but complex layers often materialize edge-wise intermediates, increasing memory traffic and limiting scalability on large graphs. We take an I/O- and arithmetic-intensity–centric view and show that widely used layers fall into three kernel families: SpMM-based convolutions, reduction-based aggregations, and attention-based layers (GATv2/Graph Transformer). For each family, we develop GPU kernels that reduce data movement, improve locality, and remain robust across realistic graphs. We also study graph reordering and find that its impact depends on the kernel mapping: it benefits neighbor-parallel (gather-dominated) kernels more consistently than feature-parallel designs. Empirically, our fused attention kernels reach up to 3.9$\times$ speedup for Graph Transformer (median 1.6$\times$), with Tensor Core (block-sparse) variants up to 7.3$\times$ on locally dense graphs; for GATv2 we reach up to 8.5$\times$ speedup (median 2.0$\times$) while reducing peak memory by up to 76$\times$ (median 6$\times$). Our degree-aware reduction kernels achieve up to 10$\times$ speedup (median 2.6$\times$). For SpMM-based layers, properly cached cuSPARSE achieves up to 8$\times$ speedup over DGL and outperforms evaluated custom baselines in the majority of evaluations. We release our implementations as drop-in replacements in our GitHub repository to support reproducible, hardware-aware GNN acceleration.} }
Endnote
%0 Conference Paper %T On Efficient Scaling of GNNs via IO-Aware Layers Implementations %A Daria Fomina %A Daniil Krasylnikov %A Alexey Boykov %A Andrey Dolgovyazov %A Vyacheslav Zhdanovskiy %A Fedor Velikonivtsev %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-fomina26a %I PMLR %P 31329--31376 %U https://proceedings.mlr.press/v306/fomina26a.html %V 306 %X Graph Neural Networks (GNNs) are bottlenecked by sparse, irregular memory access. Popular frameworks such as DGL and PyTorch Geometric support general message passing, but complex layers often materialize edge-wise intermediates, increasing memory traffic and limiting scalability on large graphs. We take an I/O- and arithmetic-intensity–centric view and show that widely used layers fall into three kernel families: SpMM-based convolutions, reduction-based aggregations, and attention-based layers (GATv2/Graph Transformer). For each family, we develop GPU kernels that reduce data movement, improve locality, and remain robust across realistic graphs. We also study graph reordering and find that its impact depends on the kernel mapping: it benefits neighbor-parallel (gather-dominated) kernels more consistently than feature-parallel designs. Empirically, our fused attention kernels reach up to 3.9$\times$ speedup for Graph Transformer (median 1.6$\times$), with Tensor Core (block-sparse) variants up to 7.3$\times$ on locally dense graphs; for GATv2 we reach up to 8.5$\times$ speedup (median 2.0$\times$) while reducing peak memory by up to 76$\times$ (median 6$\times$). Our degree-aware reduction kernels achieve up to 10$\times$ speedup (median 2.6$\times$). For SpMM-based layers, properly cached cuSPARSE achieves up to 8$\times$ speedup over DGL and outperforms evaluated custom baselines in the majority of evaluations. We release our implementations as drop-in replacements in our GitHub repository to support reproducible, hardware-aware GNN acceleration.
APA
Fomina, D., Krasylnikov, D., Boykov, A., Dolgovyazov, A., Zhdanovskiy, V. & Velikonivtsev, F.. (2026). On Efficient Scaling of GNNs via IO-Aware Layers Implementations. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:31329-31376 Available from https://proceedings.mlr.press/v306/fomina26a.html.

Related Material