The International Conference for High Performance Computing, Networking, Storage, and Analysis

Workshops Archive

A Study of Performance Portability of Low-bit Fused Matrix-Vector Multiplication Kernels in SYCL


Workshop: 12th Workshop on Accelerator Programming and Directives (WACCPD 2025)

Authors: Zheming Jin (ORNL)

Abstract: Compared to CUDA, SYCL is a portable programming model for various hardware accelerators. In this paper, we study performance portability of low-bit fused general matrix-vector multiplication kernels in SYCL on vendors’ graphics processing units (GPUs). We introduce the use case, explain the kernel implementations in details, evaluate the performance of the CUDA, HIP, and SYCL kernels on datacenter, desktop, and laptop GPUs, and investigate the causes of performance gaps. We find that loop unrolling, kernel dispatch overhead, and sum reduction contribute to the gaps. We hope that the findings provide valuable feedback for the development of the SYCL ecosystem.


Back to 12th Workshop on Accelerator Programming and Directives (WACCPD 2025) Archive Listing Back to Full Workshop Archive Listing