Presentation
A Study of Performance Portability of Low-bit Fused Matrix-Vector Multiplication Kernels in SYCL
DescriptionCompared to CUDA, SYCL is a portable programming model for
various hardware accelerators. In this paper, we study
performance portability of low-bit fused general matrix-vector
multiplication kernels in SYCL on vendors’ graphics processing
units (GPUs). We introduce the use case, explain the kernel
implementations in details, evaluate the performance of the
CUDA, HIP, and SYCL kernels on datacenter, desktop, and
laptop GPUs, and investigate the causes of performance gaps.
We find that loop unrolling, kernel dispatch overhead, and sum
reduction contribute to the gaps. We hope that the findings
provide valuable feedback for the development of the SYCL
ecosystem.
various hardware accelerators. In this paper, we study
performance portability of low-bit fused general matrix-vector
multiplication kernels in SYCL on vendors’ graphics processing
units (GPUs). We introduce the use case, explain the kernel
implementations in details, evaluate the performance of the
CUDA, HIP, and SYCL kernels on datacenter, desktop, and
laptop GPUs, and investigate the causes of performance gaps.
We find that loop unrolling, kernel dispatch overhead, and sum
reduction contribute to the gaps. We hope that the findings
provide valuable feedback for the development of the SYCL
ecosystem.
Event Type
Workshop
TimeMonday, 17 November 202511:30am - 12:00pm CST
Location266
