BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201805Z
LOCATION:266
DTSTART;TZID=America/Chicago:20251117T110000
DTEND;TZID=America/Chicago:20251117T113000
UID:submissions.supercomputing.org_SC25_sess218_ws_waccpd103@linklings.com
SUMMARY:Towards Efficient Load Balancing BFS on GPUs: One Code for AMD, In
 tel & Nvidia
DESCRIPTION:Kaan Olgu (University of Bristol), Tobias Kenter (Paderborn Un
 iversity), Jose Nunez-Yanez (Linkoping University), and Simon McIntosh-Smi
 th and Tom Deakin (University of Bristol)\n\nEfficient graph processing is
  essential for a wide range of applications.\nScalability and memory acces
 s patterns are still a challenge, especially with the Breadth-First Search
  algorithm. This work focuses on leveraging multi-GPU HPC nodes with peer-
 to-peer support of the Intel oneAPI implementation of SYCL.\nWe propose th
 ree GPU-based load-balancing methods: work-group localisation for efficien
 t data access, even workload distribution for higher GPU occupancy, and a 
 hybrid strided-access approach for heuristic balancing. These methods ensu
 re performance, portability, and productivity with a unified codebase.\nOu
 r proposed methodologies outperform state-of-the-art single-GPU implementa
 tions based on CUDA on synthetic RMAT graphs. We analysed BFS performance 
 across NVIDIA A100, Intel Max 1550, and AMD MI300X GPUs, achieving a peak 
 performance of 153.27 GTEPS on an RMAT25-64 graph using 8 GPUs on the NVID
 IA A100. Furthermore, our work handles RMAT graphs up to scale 29, achievi
 ng superior performance on synthetic graphs and competitive results on rea
 l-world datasets.\n\nRecording: Livestreamed, Recorded\n\nRegistration Cat
 egory: Technical Program Reg Pass, Workshop Reg Pass\n\nSession Chairs: An
 dreas Herten (Forschungszentrum Jülich, Jülich Supercomputing Centre (JSC)
 ); Rabab Alomairy (Massachusetts Institute of Technology (MIT), King Abdul
 lah University of Science and Technology (KAUST)); and Jorge Luis Galvez V
 allejo (Australian National University)\n\n
END:VEVENT
END:VCALENDAR
