QGTC: Accelerating Quantized Graph Neural Networks via GPU Tensor Core (PPoPP 2022 - Main Conference)

Who

Yuke Wang, Boyuan Feng, Yufei Ding

Track

PPoPP 2022 Main Conference

Time Zone

The program is currently displayed in (GMT-04:00) Eastern Time (US & Canada).

Use conference time zone: (GMT-04:00) Eastern Time (US & Canada)Select other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Mon 4 Apr 2022 12:50 - 13:05 - Session 3 Chair(s): Bin Ren

Abstract

Over the most recent years, quantized graph neural network (QGNN) attracts lots of research and industry attention due to its high robustness and low computation and memory overhead. Unfortunately, the performance gains of QGNN have never been realized on modern GPU platforms. To this end, we propose the first Tensor Core (TC) based computing framework, \textbf{QGTC}, to support any-bitwidth computation for QGNNs on GPUs. We introduce a novel quantized low-bit arithmetic design based on the low-bit data representation and bit-decomposed computation. We craft a novel TC-tailored CUDA kernel design by incorporating 3D-stacked bit compression, zero-tile jumping, and non-zero tile reuse technique to improve the performance systematically. We incorporate an effective bandwidth-optimized subgraph packing strategy to maximize the transferring efficiency between CPU host and GPU device. We integrate QGTC with PyTorch for better programmability and extensibility. Extensive experiments demonstrate that QGTC can achieve evident inference speedup (on average $2.7\times$) compared with the state-of-the-art DGL framework across diverse settings.

Yuke Wang

UC Santa Barbara

United States

Boyuan Feng

University of California Santa Barbara

Yufei Ding

University of California at Santa Barbara