Global Neighbor Sampling for Mixed CPU-GPU Training on Giant Graphs

Jialin Dong; Da Zheng; Lin F. Yang; George Karypis

doi:10.1145/3447548.3467437

Global Neighbor Sampling for Mixed CPU-GPU Training on Giant Graphs

Jialin Dong, Da Zheng, Lin F. Yang, George Karypis

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution

10 Scopus citations

Abstract

Graph neural networks (GNNs) are powerful tools for learning from graph data and are widely used in various applications such as social network recommendation, fraud detection, and graph search. The graphs in these applications are typically large, usually containing hundreds of millions of nodes. Training GNN models on such large graphs efficiently remains a big challenge. Despite a number of sampling-based methods have been proposed to enable mini-batch training on large graphs, these methods have not been proved to work on truly industry-scale graphs, which require GPUs or mixed CPU-GPU training. The state-of-the-art sampling-based methods are usually not optimized for these real-world hardware setups, in which data movement between CPUs and GPUs is a bottleneck. To address this issue, we propose Global Neighborhood Sampling that aims at training GNNs on giant graphs specifically for mixed CPU-GPU training. The algorithm samples a global cache of nodes periodically for all mini-batches and stores them in GPUs. This global cache allows in-GPU importance sampling of mini-batches, which drastically reduces the number of nodes in a mini-batch, especially in the input layer, to reduce data copy between CPU and GPU and mini-batch computation without compromising the training convergence rate or model accuracy. We provide a highly efficient implementation of this method and show that our implementation outperforms an efficient node-wise neighbor sampling baseline by a factor of 2× ∼ 4× on giant graphs. It outperforms an efficient implementation of LADIES with small layers by a factor of 2× ∼ 14× while achieving much higher accuracy than LADIES. We also theoretically analyze the proposed algorithm and show that with cached node data of a proper size, it enjoys a comparable convergence rate as the underlying node-wise sampling method.

Original language	English (US)
Title of host publication	KDD 2021 - Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
Publisher	Association for Computing Machinery
Pages	289-299
Number of pages	11
ISBN (Electronic)	9781450383325
DOIs	https://doi.org/10.1145/3447548.3467437
State	Published - Aug 14 2021
Externally published	Yes
Event	27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2021 - Virtual, Online, Singapore Duration: Aug 14 2021 → Aug 18 2021

Publication series

Name	Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining

Conference

Conference	27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2021
Country/Territory	Singapore
City	Virtual, Online
Period	8/14/21 → 8/18/21

Bibliographical note

Publisher Copyright:
© 2021 Owner/Author.

Keywords

graph neural networks
mixed CPU-GPU training
neighbor sampling

Access

10.1145/3447548.3467437

OpenUrl availability

Full text

Cite this

Dong, J., Zheng, D., Yang, L. F., & Karypis, G. (2021). Global Neighbor Sampling for Mixed CPU-GPU Training on Giant Graphs. In KDD 2021 - Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 289-299). (Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining). Association for Computing Machinery. https://doi.org/10.1145/3447548.3467437

Global Neighbor Sampling for Mixed CPU-GPU Training on Giant Graphs. / Dong, Jialin; Zheng, Da; Yang, Lin F. et al.
KDD 2021 - Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery, 2021. p. 289-299 (Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining).

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution

Dong, J, Zheng, D, Yang, LF & Karypis, G 2021, Global Neighbor Sampling for Mixed CPU-GPU Training on Giant Graphs. in KDD 2021 - Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Association for Computing Machinery, pp. 289-299, 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2021, Virtual, Online, Singapore, 8/14/21. https://doi.org/10.1145/3447548.3467437

Dong J, Zheng D, Yang LF, Karypis G. Global Neighbor Sampling for Mixed CPU-GPU Training on Giant Graphs. In KDD 2021 - Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery. 2021. p. 289-299. (Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining). doi: 10.1145/3447548.3467437

@inproceedings{c921062f21fd4fca81e9c6783654ebe7,

title = "Global Neighbor Sampling for Mixed CPU-GPU Training on Giant Graphs",

abstract = "Graph neural networks (GNNs) are powerful tools for learning from graph data and are widely used in various applications such as social network recommendation, fraud detection, and graph search. The graphs in these applications are typically large, usually containing hundreds of millions of nodes. Training GNN models on such large graphs efficiently remains a big challenge. Despite a number of sampling-based methods have been proposed to enable mini-batch training on large graphs, these methods have not been proved to work on truly industry-scale graphs, which require GPUs or mixed CPU-GPU training. The state-of-the-art sampling-based methods are usually not optimized for these real-world hardware setups, in which data movement between CPUs and GPUs is a bottleneck. To address this issue, we propose Global Neighborhood Sampling that aims at training GNNs on giant graphs specifically for mixed CPU-GPU training. The algorithm samples a global cache of nodes periodically for all mini-batches and stores them in GPUs. This global cache allows in-GPU importance sampling of mini-batches, which drastically reduces the number of nodes in a mini-batch, especially in the input layer, to reduce data copy between CPU and GPU and mini-batch computation without compromising the training convergence rate or model accuracy. We provide a highly efficient implementation of this method and show that our implementation outperforms an efficient node-wise neighbor sampling baseline by a factor of 2× ∼ 4× on giant graphs. It outperforms an efficient implementation of LADIES with small layers by a factor of 2× ∼ 14× while achieving much higher accuracy than LADIES. We also theoretically analyze the proposed algorithm and show that with cached node data of a proper size, it enjoys a comparable convergence rate as the underlying node-wise sampling method.",

keywords = "graph neural networks, mixed CPU-GPU training, neighbor sampling",

author = "Jialin Dong and Da Zheng and Yang, {Lin F.} and George Karypis",

note = "Publisher Copyright: {\textcopyright} 2021 Owner/Author.; 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2021 ; Conference date: 14-08-2021 Through 18-08-2021",

year = "2021",

month = aug,

day = "14",

doi = "10.1145/3447548.3467437",

language = "English (US)",

series = "Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining",

publisher = "Association for Computing Machinery",

pages = "289--299",

booktitle = "KDD 2021 - Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining",

}

TY - GEN

T1 - Global Neighbor Sampling for Mixed CPU-GPU Training on Giant Graphs

AU - Dong, Jialin

AU - Zheng, Da

AU - Yang, Lin F.

AU - Karypis, George

PY - 2021/8/14

Y1 - 2021/8/14

N2 - Graph neural networks (GNNs) are powerful tools for learning from graph data and are widely used in various applications such as social network recommendation, fraud detection, and graph search. The graphs in these applications are typically large, usually containing hundreds of millions of nodes. Training GNN models on such large graphs efficiently remains a big challenge. Despite a number of sampling-based methods have been proposed to enable mini-batch training on large graphs, these methods have not been proved to work on truly industry-scale graphs, which require GPUs or mixed CPU-GPU training. The state-of-the-art sampling-based methods are usually not optimized for these real-world hardware setups, in which data movement between CPUs and GPUs is a bottleneck. To address this issue, we propose Global Neighborhood Sampling that aims at training GNNs on giant graphs specifically for mixed CPU-GPU training. The algorithm samples a global cache of nodes periodically for all mini-batches and stores them in GPUs. This global cache allows in-GPU importance sampling of mini-batches, which drastically reduces the number of nodes in a mini-batch, especially in the input layer, to reduce data copy between CPU and GPU and mini-batch computation without compromising the training convergence rate or model accuracy. We provide a highly efficient implementation of this method and show that our implementation outperforms an efficient node-wise neighbor sampling baseline by a factor of 2× ∼ 4× on giant graphs. It outperforms an efficient implementation of LADIES with small layers by a factor of 2× ∼ 14× while achieving much higher accuracy than LADIES. We also theoretically analyze the proposed algorithm and show that with cached node data of a proper size, it enjoys a comparable convergence rate as the underlying node-wise sampling method.

AB - Graph neural networks (GNNs) are powerful tools for learning from graph data and are widely used in various applications such as social network recommendation, fraud detection, and graph search. The graphs in these applications are typically large, usually containing hundreds of millions of nodes. Training GNN models on such large graphs efficiently remains a big challenge. Despite a number of sampling-based methods have been proposed to enable mini-batch training on large graphs, these methods have not been proved to work on truly industry-scale graphs, which require GPUs or mixed CPU-GPU training. The state-of-the-art sampling-based methods are usually not optimized for these real-world hardware setups, in which data movement between CPUs and GPUs is a bottleneck. To address this issue, we propose Global Neighborhood Sampling that aims at training GNNs on giant graphs specifically for mixed CPU-GPU training. The algorithm samples a global cache of nodes periodically for all mini-batches and stores them in GPUs. This global cache allows in-GPU importance sampling of mini-batches, which drastically reduces the number of nodes in a mini-batch, especially in the input layer, to reduce data copy between CPU and GPU and mini-batch computation without compromising the training convergence rate or model accuracy. We provide a highly efficient implementation of this method and show that our implementation outperforms an efficient node-wise neighbor sampling baseline by a factor of 2× ∼ 4× on giant graphs. It outperforms an efficient implementation of LADIES with small layers by a factor of 2× ∼ 14× while achieving much higher accuracy than LADIES. We also theoretically analyze the proposed algorithm and show that with cached node data of a proper size, it enjoys a comparable convergence rate as the underlying node-wise sampling method.

KW - graph neural networks

KW - mixed CPU-GPU training

KW - neighbor sampling

UR - http://www.scopus.com/inward/record.url?scp=85114931105&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=85114931105&partnerID=8YFLogxK

U2 - 10.1145/3447548.3467437

DO - 10.1145/3447548.3467437

M3 - Conference contribution

AN - SCOPUS:85114931105

T3 - Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining

SP - 289

EP - 299

BT - KDD 2021 - Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining

PB - Association for Computing Machinery

T2 - 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2021

Y2 - 14 August 2021 through 18 August 2021

ER -

Global Neighbor Sampling for Mixed CPU-GPU Training on Giant Graphs

Abstract

Publication series

Conference

Bibliographical note

Keywords

Access

OpenUrl availability

Other files and links

Fingerprint

Cite this