Optimization of Large-Scale Graph Processing with Advanced Partitioning and Thread Grouping

Authors

  • Muhammad Numan University of Peshawar
  • Abdul Haseeb Malik University of Peshawar
  • Qazi Ejaz Ali University of Peshawar
  • Waheed Ur Rehman University of Peshawar
  • Muhammad Haseeb University of Peshawar

Keywords:

Graph Partitioning, Large-Scale Graphs, Distributed Computing, GPU Acceleration

Abstract

Large-scale graphs are a natural way of modeling entities with complex relationships. Irregular structures and relationships complicate the processing of these large-scale graphs. Partitioning enables efficient processing of these graphs; however, it may result in information loss and imbalanced workload distribution. The processing of large graphs is greatly limited by these challenges, as they make it difficult to scale and efficiently process them. This paper provides a large-scale graph processing framework that involves graph partitioning and thread grouping to enhance memory usage and computational efficiency of a GPU-based system. The proposed partitioning algorithm creates subgraphs with minimum cross-partitions and minimum cross-partition edges. A thread grouping algorithm efficiently processes these partitions in parallel, minimizing inter-partition connectivity. It creates an execution pack of partitions that can be efficiently processed on a GPU at runtime based on available memory, while allowing hardware resources to be used elastically. The GraphSAGE architecture is used to evaluate the framework with several real-world benchmark datasets, such as OGBN-Products, Reddit, and LiveJournal. The experimental findings indicate that the proposed execution strategy achieves a substantial decrease in memory pressure and preserves competitive classification accuracy. The proposed work provides better stability and scalability of training without compromising the performance of the model, which is well-suited for the efficient processing of large-scale graphs.

References

D. Salwasser, D. Seemaier, L. Gottesbüren, and P. Sanders, “Tera-Scale Multilevel Graph Partitioning,” Oct. 2024, Accessed: Jul. 27, 2026. [Online]. Available: https://arxiv.org/pdf/2410.19119

O. Omolayo, O. Oloruntoba, S. Adepoju, and K. Audu, “Comparative Analysis of Graph Partitioning Strategies for Enhancing Scalability and Performance in Distributed Graph Databases,” Jun. 2025, doi: 10.21203/RS. 3. RS-6900355/V1.

X. Cheng et al., "LPS-GNN: Deploying Graph Neural Networks on Graphs with 100 Billion Edges,” ACM Trans. Knowl. Discov. Data, vol. 20, no. 5, Jul. 2025, doi: 10.1145/3801100.

H. Gao, X. Liao, Z. Shao, K. Li, J. Chen, and H. Jin, “A survey on dynamic graph processing on GPUs: concepts, terminologies, and systems,” Front. Comput. Sci. 2024 184, vol. 18, no. 4, pp. 184106-, Dec. 2023, doi: 10.1007/S11704-023-2656-1.

Z. Cai et al., “DSP: Efficient GNN Training with Multiple GPUs,” Proc. ACM SIGPLAN Symp. Princ. Pract. Parallel Program. PPOPP, vol. 1, pp. 392–404, Feb. 2023, doi: 10.1145/3572848.3577528; PAGE:STRING:ARTICLE/CHAPTER.

L. Gottesbüren, T. Heuer, P. Sanders, C. Schulz, and D. Seemaier, “Deep Multilevel Graph Partitioning,” Leibniz Int. Proc. Informatics, LIPIcs, vol. 204, May 2021, doi: 10.4230/LIPIcs.ESA.2021.48.

M. S. Gilbert, K. Madduri, E. G. Boman, and S. Rajamanickam, “Jet: Multilevel Graph Partitioning on Graphics Processing Units,” Apr. 2023, Accessed: Jul. 27, 2026. [Online]. Available: https://arxiv.org/pdf/2304.13194

R. Waleffe, D. Sarda, J. Mohoney, E.-V. Vlatakis-Gkaragkounis, T. Rekatsinas, and S. Venkataraman, “Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks,” Feb. 2025, Accessed: Jul. 27, 2026. [Online]. Available: https://arxiv.org/pdf/2502.17846

H. Cui, F. Cao, and R. Liu, “A multi-objective partitioning algorithm for large-scale graph based on NSGA-II,” Expert Syst. Appl., vol. 263, p. 125756, Mar. 2025, doi: 10.1016/J.ESWA.2024.125756.

Z. Ye et al., “Deep Learning Workload Scheduling in GPU Datacenters: A Survey,” ACM Comput. Surv., vol. 56, no. 6, p. 146, Jun. 2024, doi: 10.1145/3638757; Journal: Csur; Wgroup: String: Acm.

“(PDF) Parmetis: Parallel graph partitioning and sparse matrix ordering library.” Accessed: Jul. 27, 2026. [Online]. Available: https://www.researchgate.net/publication/238705993_Parmetis_Parallel_graph_partitioning_and_sparse_matrix_ordering_library

A. Gupta, “Fast and effective algorithms for graph partitioning and sparse-matrix ordering,” IBM J. Res. Dev., vol. 41, no. 1–2, pp. 171–183, 1997, doi: 10.1147/RD.411.0171.

A. Pirova, I. Meyerov, E. Kozinov, and S. Lebedev, “PMORSy: parallel sparse matrix ordering software for fill-in minimization†,” Optim. Methods Softw., vol. 32, no. 2, pp. 274–289, Mar. 2017, doi: 10.1080/10556788.2016.1193177;REQUESTEDJOURNAL:Journal:Opms;Wgroup:String:Acm.

L. Meng et al., “A Survey of Distributed Graph Algorithms on Massive Graphs,” ACM Comput. Surv., vol. 57, no. 2, p. 39, Oct. 2024, doi: 10.1145/3694966;Requestedjournal: Csur;Taxonomy:Taxonomy:Acm-Pubtype;Pagegroup:String:Publication.

D. Tripathy, A. Abdolrashidi, L. N. Bhuyan, L. Zhou, and D. Wong, “PAVER: Locality Graph-Based Thread Block Scheduling for GPUs,” ACM Trans. Archit. Code Optim., vol. 18, no. 3, Jun. 2021, doi: 10.1145/3451164; Page: String: Article/Chapter.

Y. Lü et al., “GraphPEG: Accelerating Graph Processing on GPUs,” ACM Trans. Archit. Code Optim., vol. 18, no. 3, Jun. 2021, doi: 10.1145/3450440; Requested journal: Taco; Serial topic: Topic: ACM-PubType.

A. Sahebi, M. Barbone, M. Procaccini, W. Luk, G. Gaydadjiev, and R. Giorgi, “Distributed large-scale graph processing on FPGAs,” J. Big Data 2023 101, vol. 10, no. 1, pp. 95-, Jun. 2023, doi: 10.1186/S40537-023-00756-X.

Y. Wang, C. Colley, B. Wheatman, J. Su, D. F. Gleich, and A. A. Chien, “How Fast Can Graph Computations Go on Fine-grained Parallel Architectures,” Jul. 2025, Accessed: Jul. 27, 2026. [Online]. Available: https://arxiv.org/pdf/2507.00949

J. He, M. Jia, Y. Chen, Z. Liu, and D. Li, “GTSM: A multi-edge-centric temporal subgraph matching framework on GPUs,” ACM Trans. Archit. Code Optim., vol. 22, no. 4, Dec. 2025, doi: 10.1145/3771286; Website: Dl-Site; Taxonomy: Acm-Pubtype; Pagegroup: String: Publication.

P. Cui, H. Liu, B. Tang, and Y. Yuan, “CGgraph: An Ultra-fast Graph Processing System on Modern Commodity CPU-GPU Co-processor,” Proc. VLDB Endow., vol. 17, no. 6, pp. 1405–1417, Feb. 2024, doi: 10.14778/3648160.3648179; Wgroup: String: Acm.

Q. Wang et al., “Efficient Graph Data Access for Out-of-Memory GPU Streaming Graph Processing,” Proc. VLDB Endow., vol. 18, no. 11, pp. 3854–3867, Jul. 2025, doi: 10.14778/3749646.3749659; Serial topic: Topic: Acm-Pubtype.

Y. Sun et al., “CGCGraph: Efficient CPU-GPU Co-execution for Concurrent Dynamic Graph Processing,” ACM Trans. Archit. Code Optim., vol. 22, no. 3, Sep. 2025, doi: 10.1145/3744904; Journal: TACO

Y. Zhang et al., “LargeGraph: An Efficient Dependency-Aware GPU-Accelerated Large-Scale Graph Processing,” ACM Trans. Archit. Code Optim., vol. 18, no. 4, p. 58, Dec. 2021, doi: 10.1145/3477603; Requested journal: TACO; serial topic: Acm-Pubtype.

S. Song et al., “Large-Scale Dynamic Graph Processing with Graphic Processing Unit-Accelerated Priority-Driven Differential Scheduling and Operation Reduction,” Appl. Sci. 2025, Vol. 15, Page 3172, vol. 15, no. 6, p. 3172, Mar. 2025, doi: 10.3390/APP15063172.

J. Zhao et al., “An Efficient ReRAM-based Accelerator for Asynchronous Iterative Graph Processing,” ACM Trans. Archit. Code Optim., vol. 22, no. 1, p. 16, Mar. 2025, doi: 10.1145/3689335; Website: Dl-Site; Taxonomy: Acm-Pubtype; Pagegroup: String:

Publication.

R. Chen, X. Weng, B. He, and M. Yang, “Large graph processing in the cloud,” Proc. ACM SIGMOD Int. Conf. Manag. Data, pp. 1123–1126, 2010, doi: 10.1145/1807167.1807297; Topic: Conference-Collections.

“GraphChi | Proceedings of the 10th USENIX conference on Operating Systems Design and Implementation.” Accessed: Jul. 27, 2026. [Online]. Available: https://dl.acm.org/doi/10.5555/2387880.2387884

S. Salihoglu and J. Widom, “GPS: A graph processing system,” ACM Int. Conf. Proceeding Ser., Jul. 2013, doi: 10.1145/2484838.2484843; Topic: Conference-Collections.

J. Shun and G. E. Blelloch, “Ligra: A lightweight graph processing framework for shared memory,” ACM SIGPLAN Not., vol. 48, no. 8, pp. 135–146, Aug. 2013, doi: 10.1145/2517327.2442530; Serial topic: Acm-Pubtype.

J. Ahn, S. Hong, S. Yoo, O. Mutlu, and K. Choi, “A scalable processing-in-memory accelerator for parallel graph processing,” Proc. - Int. Symp. Comput. Archit., vol. 13-17-June-2015, pp. 105–117, Jun. 2015, doi: 10.1145/2749469.2750386.

“Gunrock: a high-performance graph processing library on the GPU: ACM SIGPLAN Notices: Vol 51, No 8.” Accessed: Jul. 27, 2026. [Online]. Available: https://dl.acm.org/doi/10.1145/3016078.2851145

T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” 5th Int. Conf. Learn. Represent. ICLR 2017 - Conf. Track Proc., 2017.

J. L. William L. Hamilton, Rex Ying, “Inductive Representation Learning on Large Graphs,” arXiv:1706.02216, 2017, doi: https://doi.org/10.48550/arXiv.1706.02216.

P. Veličković, A. Casanova, P. Liò, G. Cucurull, A. Romero, and Y. Bengio, “Graph Attention Networks,” 6th Int. Conf. Learn. Represent. ICLR 2018 - Conf. Track Proc., Oct. 2017, doi: 10.1007/978-3-031-01587-8_7.

K. Xu, S. Jegelka, W. Hu, and J. Leskovec, “How Powerful are Graph Neural Networks?,” 7th Int. Conf. Learn. Represent. ICLR 2019, Oct. 2018. Accessed: Jul. 27, 2026. [Online]. Available: https://arxiv.org/pdf/1810.00826

W.-L. Chiang, X. Liu, S. Si, Y. Li, S. Bengio, and C.-J. Hsieh, “Cluster-GCN: An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks,” Proc. ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., pp. 257–266, Aug. 2019, doi: 10.1145/3292500.3330925.

Downloads

Published

2026-07-18

How to Cite

Muhammad Numan, Abdul Haseeb Malik, Qazi Ejaz Ali, Waheed Ur Rehman, & Muhammad Haseeb. (2026). Optimization of Large-Scale Graph Processing with Advanced Partitioning and Thread Grouping. International Journal of Innovations in Science & Technology, 8(4), 1639–1659. Retrieved from https://journal.50sea.com/index.php/IJIST/article/view/1944