BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201435Z
LOCATION:266
DTSTART;TZID=America/Chicago:20251117T090000
DTEND;TZID=America/Chicago:20251117T173000
UID:submissions.supercomputing.org_SC25_sess218@linklings.com
SUMMARY:12th Workshop on Accelerator Programming and Directives (WACCPD 20
 25)
DESCRIPTION:Heterogeneous node architectures are ubiquitous in today’s HPC
  landscape. Exploiting the compute capability, while maintaining code port
 ability and maintainability, necessitates effective accelerator programmin
 g approaches. The use of these programming approaches remains a research a
 ctivity, and there are many possible trade-offs between performance, porta
 bility, maintainability, and ease of use that must be considered. Addition
 ally, new heterogeneous computing concepts are being deployed, like ML/AI 
 chips and QPUs, introducing challenges related to algorithms, portability,
  and standardization of programming models. The WACCPD workshop highlights
  the improvements over state-of-the-art through accepted papers and talks.
  The event will also foster discussion with invited talks and a panel to d
 raw the community’s attention to key areas that will facilitate the transi
 tion to accelerator-based HPC, including AI, and quantum computing. The wo
 rkshop aims to showcase all aspects of innovative language features, lesso
 ns learned while using directives/abstractions to migrate scientific code,
  and experiences using novel accelerator architectures, among others.\n\nL
 unch break (on your own)\n---------------------\nAfternoon Break - Worksho
 p on Accelerator Programming and Directives (WACCPD 2025)\n---------------
 ------\nTowards Efficient Load Balancing BFS on GPUs: One Code for AMD, In
 tel & Nvidia\n\nEfficient graph processing is essential for a wide range o
 f applications.\nScalability and memory access patterns are still a challe
 nge, especially with the Breadth-First Search algorithm. This work focuses
  on leveraging multi-GPU HPC nodes with peer-to-peer support of the Intel 
 oneAPI implementation...\n\n\nKaan Olgu (University of Bristol), Tobias Ke
 nter (Paderborn University), Jose Nunez-Yanez (Linkoping University), and 
 Simon McIntosh-Smith and Tom Deakin (University of Bristol)\n-------------
 --------\nScalable Neural Network Training: Distributed Data-Parallel Appr
 oaches\n\nTraining large neural networks is computationally demanding and 
 often limited by synchronization overhead in distributed environments. Tra
 ditional data-parallel frameworks, such as Horovod or PyTorch DDP, average
  gradients at every batch, which can limit scalability due to communicatio
 n bottlenecks....\n\n\nFernando Vazquez-Novoa (Barcelona Supercomputing Ce
 nter (BSC)), Pedro López and José Flich (Universidad Politecnica de Valenc
 ia), and Rosa M. Badia (Barcelona Supercomputing Center (BSC))\n----------
 -----------\nMorning Break - Workshop on Accelerator Programming and Direc
 tives (WACCPD 2025)\n---------------------\nInvited Talk: Weather and Clim
 ate Codes in the AI Era\n\nWe will start the discussion with a landmark ac
 hievement: the first global simulation of the full Earth system at a 1.25 
 km grid spacing. Our talk will focus on how we used the Alps supercomputer
  to model the intricate flow of energy, water, and carbon across the atmos
 phere, ocean, and land. We will...\n\n\nTorsten Hoefler (ETH Zürich, Swiss
  National Supercomputing Centre (CSCS))\n---------------------\nPorting a 
 Fortran plasma simulation to Exascale on AMD GPUs using both OpenMP and Ko
 kkos\n\nThis paper presents the 2-step work undertaken to port GYSELA, a p
 etascale Fortran simulation code for turbulence in tokamak plasmas, to GPU
 s. The initial porting process using OpenMP offloading allowed for good pe
 rformance in most of the code, with the exception of the collision operato
 r, which bec...\n\n\nEtienne Malaboeuf (CINES, CEA); Mathieu Peybernes (EP
 FL, SCITAS); Kévin Obrejan and Julien Julien (CEA); Emily Bourne (SCITAS, 
 EPFL); and Virginie Grandgirard (CEA)\n---------------------\nBest Paper A
 ward & Closing\n---------------------\nTwelfth Workshop on Accelerator Pro
 gramming and Directives (WACCPD 2025)\n\nHeterogeneous node architectures 
 are ubiquitous in today’s HPC landscape. Exploiting the compute capability
 , while maintaining code portability and maintainability, necessitates eff
 ective accelerator programming approaches. The use of these programming ap
 proaches remains a research activity, a...\n\n\nAndreas Herten (Forschungs
 zentrum Jülich, Jülich Supercomputing Centre (JSC)); Rabab Alomairy (Massa
 chusetts Institute of Technology (MIT)); and Jorge Luis Gálvez Vallejo (Au
 stralian National University)\n---------------------\nMojo: MLIR-based Per
 formance-Portable HPC Science Kernels on GPUs for the Python Ecosystem\n\n
 We explore the performance and portability of the novel Mojo language for 
 scientific computing workloads on GPUs. As the first language based on the
  LLVM's Multi-Level Intermediate Representation (MLIR) compiler infrastruc
 ture, Mojo aims to close performance and productivity gaps by combining Py
 thon...\n\n\nWilliam Godoy (Oak Ridge National Laboratory (ORNL)); Tatiana
  Melnichenko (University of Tennessee, Knoxville; Oak Ridge National Labor
 atory (ORNL)); and Pedro Valero-Lara, Wael Elwasif, Philip Fackler, Rafael
  Ferreira Da Silva, Keita Teranishi, and Jeffrey Vetter (Oak Ridge Nationa
 l Laboratory (ORNL))\n---------------------\nPanel Discussion\n\nAndreas H
 erten (Jülich Supercomputing Centre); Tatiana Melnichenko (University of T
 ennessee, Knoxville; Oak Ridge National Laboratory (ORNL)); Kaan Olgu (Uni
 versity of Bristol); Zheming Jin (Oak Ridge National Laboratory); Etienne 
 Malaboeuf, Michele Martinelli, and Richard Schulze; and Fernando Vazquez-N
 ovoa (Barcelona Supercomputing Center)\n---------------------\nInvited Tal
 k: Bridging Emerging Accelerator Architectures and HPC Production Systems:
  From Sandia Testbeds to Performance-Portable Software\n\nSandia National 
 Laboratories is charting a full-stack path that takes experimental acceler
 ator hardware from prototype testbeds into the high-performance production
  systems running our mission codes. In the first part of this talk, I’ll s
 urvey our hardware prototyping journey—from the Ad...\n\n\nSimon Garcia de
  Gonzalo (Sandia National Laboratories)\n---------------------\nA Study of
  Performance Portability of Low-bit Fused Matrix-Vector Multiplication Ker
 nels in SYCL\n\nCompared to CUDA, SYCL is a portable programming model for
 \nvarious hardware accelerators. In this paper, we study\nperformance port
 ability of low-bit fused general matrix-vector\nmultiplication kernels in 
 SYCL on vendors’ graphics processing\nunits (GPUs). We introduce the use c
 ase, explain the k...\n\n\nZheming Jin (ORNL)\n---------------------\nRedu
 ction-Aware Directive-Based Programming via Multi-Dimensional Homomorphism
 s\n\nDirective-based programming is a productive way to target parallel ar
 chitectures like GPUs and CPUs. Popular solutions such as OpenMP and OpenA
 CC are widely used because they are simple and broadly applicable to gener
 al-purpose codebases. However, they often fail to deliver consistently hig
 h and por...\n\n\nRichard Schulze, Sergei Gorlatch, and Ari Rasch (Univers
 ity of Muenster)\n---------------------\nBridging FPGA and GPU over PCIe: 
 A Low-Latency Communication Path using AVX-512\n\nWe introduce a communica
 tion mechanism bridging accelerators like GPUs and PCIe-based FPGA devices
  using Programmed I/O as an alternative to Direct Memory Access data trans
 missions: less than 2 microseconds one-way latency for small message trans
 fers is achieved when the FPGA operates as Network Int...\n\n\nMichele Mar
 tinelli (National Institute for Nuclear Physics (INFN)); Carlotta Chiarini
  (National Institute for Nuclear Physics, Sapienza University of Rome); An
 drea Biagioni (National Institute for Nuclear Physics); Paolo Cretaro (Nat
 ional Institute for Nuclear Physics (Currently Unaffiliated)); and Ottorin
 o Frezza, Francesca Lo Cicero, Alessandro Lonardo, Pierpaolo Perticaroli, 
 Francesco Simula, Luca Pontisso, Cristian Rossi, and Piero Vicini (Nationa
 l Institute for Nuclear Physics)\n\nRecording: Livestreamed, Recorded\n\nR
 egistration Category: Technical Program Reg Pass, Workshop Reg Pass\n\nSes
 sion Chairs: Andreas Herten (Forschungszentrum Jülich, Jülich Supercomputi
 ng Centre (JSC)); Rabab Alomairy (Massachusetts Institute of Technology (M
 IT), King Abdullah University of Science and Technology (KAUST)); and Jorg
 e Luis Galvez Vallejo (Australian National University)
END:VEVENT
END:VCALENDAR
