BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201248Z
LOCATION:Second Floor Atrium
DTSTART;TZID=America/Chicago:20251120T080000
DTEND;TZID=America/Chicago:20251120T170000
UID:submissions.supercomputing.org_SC25_sess533_post282@linklings.com
SUMMARY:Optimizing Collectives with Large Payloads on GPU-Based Supercompu
 ters
DESCRIPTION:Siddharth Singh (NVIDIA Corporation, University of Maryland); 
 Mahua Singh (IIT Guwahati); and Keshav Pradeep and Abhinav Bhatele (Univer
 sity of Maryland)\n\nWe evaluate the current state of collective communica
 tion on GPU-based supercomputers for large language model (LLM) training a
 t scale. Existing libraries such as RCCL and Cray-MPICH exhibit critical l
 imitations on systems such as Frontier—Cray-MPICH underutilizes network an
 d compute resources, while RCCL suffers from severe scalability issues. To
  address these challenges, we introduce PCCL, a communication library with
  highly optimized implementations of all-gather and reduce-scatter operati
 ons tailored for distributed deep learning workloads. PCCL is designed to 
 maximally utilize all available network and compute resources and to scale
  efficiently to thousands of GPUs. It achieves substantial performance imp
 rovements, delivering 6-33x speedups over RCCL and 28-70x over Cray-MPICH 
 for all-gather on 2,048 GCDs of Frontier. These gains translate directly t
 o end-to-end performance: in large-scale GPT-3-style training, PCCL provid
 es up to 60% and 40% speedups over RCCL for 7B and 13B parameter models, r
 espectively.\n\nTag: Research & ACM SRC Posters\n\nRegistration Category: 
 Technical Program Reg Pass\n\nSession Chairs: Kento Sato (RIKEN Center for
  Computational Science (R-CCS)); Chris Schlipalius (Pawsey Supercomputing 
 Research Centre; Commonwealth Scientific and Industrial Research Organisat
 ion (CSIRO), Australia); and Anja Gerbes (Georg-August-Universität Götting
 en)\n\n
END:VEVENT
END:VCALENDAR
