BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201803Z
LOCATION:266
DTSTART;TZID=America/Chicago:20251117T113000
DTEND;TZID=America/Chicago:20251117T120000
UID:submissions.supercomputing.org_SC25_sess218_ws_waccpd108@linklings.com
SUMMARY:A Study of Performance Portability of Low-bit Fused Matrix-Vector 
 Multiplication Kernels in SYCL
DESCRIPTION:Zheming Jin (ORNL)\n\nCompared to CUDA, SYCL is a portable pro
 gramming model for\nvarious hardware accelerators. In this paper, we study
 \nperformance portability of low-bit fused general matrix-vector\nmultipli
 cation kernels in SYCL on vendors’ graphics processing\nunits (GPUs). We i
 ntroduce the use case, explain the kernel\nimplementations in details, eva
 luate the performance of the\nCUDA, HIP, and SYCL kernels on datacenter, d
 esktop, and\nlaptop GPUs, and investigate the causes of performance gaps.\
 nWe find that loop unrolling, kernel dispatch overhead, and sum\nreduction
  contribute to the gaps. We hope that the findings\nprovide valuable feedb
 ack for the development of the SYCL\necosystem.\n\nRecording: Livestreamed
 , Recorded\n\nRegistration Category: Technical Program Reg Pass, Workshop 
 Reg Pass\n\nSession Chairs: Andreas Herten (Forschungszentrum Jülich, Jüli
 ch Supercomputing Centre (JSC)); Rabab Alomairy (Massachusetts Institute o
 f Technology (MIT), King Abdullah University of Science and Technology (KA
 UST)); and Jorge Luis Galvez Vallejo (Australian National University)\n\n
END:VEVENT
END:VCALENDAR
