Extending the SYCL Joint Matrix for Binarized Neural Networks

Zheming Jin · 2024

In contrast to the warp matrix-multiplication application-programming interface (WMMA) for tensor hardware programming in Compute Unified Device Architecture (CUDA), the SYCL joint matrix extension is an experimental feature for devices that contain a matrix hardware. While the WMMA defines 1-bit precision and bit arithmetic operations as an experimental feature for applications such as binarized neural networks (BNNs), the SYCL joint matrix extension lacks such support. This paper describes the extension to the SYCL matrix interface to leverage bit-capability of the tensor hardware. Then, we share the experience of migrating BNNs from CUDA to SYCL. Finally, we evaluate the performance of BNNs in CUDA and SYCL on graphics processing units. We hope that the results will be useful for improving portability of the SYCL programming model.

Read the paper · More papers on PaperTik