BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201259Z
LOCATION:Second Floor Atrium
DTSTART;TZID=America/Chicago:20251121T080000
DTEND;TZID=America/Chicago:20251121T120000
UID:submissions.supercomputing.org_SC25_sess620_post253@linklings.com
SUMMARY:Understanding Communication Bottlenecks in Multi-Node LLM Inferenc
 e
DESCRIPTION:Prajwal Singhania (University of Maryland); Siddharth Singh (U
 niversity of Maryland, NVIDIA Corporation); Lannie Dalton Hough and Ishan 
 Revankar (University of Maryland); Harshitha Menon and Charles Jekel (Lawr
 ence Livermore National Laboratory (LLNL)); and Abhinav Bhatele (Universit
 y of Maryland)\n\nAs large language models (LLMs) grow in parameter count,
  efficient generation requires inference to scale beyond a single node. Cu
 rrent approaches use tensor parallelism (TP) or pipeline parallelism (PP),
  but TP incurs high communication volume, while PP suffers from pipeline b
 ubbles and is unsuitable for latency-critical scenarios. We present Yalis 
 (Yet Another LLM Inference System), a lightweight and modular distributed 
 inference framework that performs comparably to existing state-of-the-art 
 systems for offline inference, while enabling rapid prototyping. Using Yal
 is, we study strong scaling of LLM inference on the Alps and Perlmutter su
 percomputers, revealing the poor scaling performance of existing paralleli
 sm strategies due to high communication overheads. We further compare the 
 all-reduce performance of NCCL and MPI in the small-message regime, findin
 g that while NCCL is efficient intra-node, MPI can outperform it cross-nod
 e for messages between 256-1024 KB. These results motivate the need for co
 mmunication-efficient parallelism strategies for multi-node LLM inference.
 \n\nTag: Research & ACM SRC Posters\n\nRegistration Category: Technical Pr
 ogram Reg Pass\n\nSession Chairs: Kento Sato (RIKEN Center for Computation
 al Science (R-CCS)); Anja Gerbes (Georg-August-Universität Göttingen); and
  Chris Schlipalius (Pawsey Supercomputing Research Centre; Commonwealth Sc
 ientific and Industrial Research Organisation (CSIRO), Australia)\n\n
END:VEVENT
END:VCALENDAR
