BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201805Z
LOCATION:261-262-265-266
DTSTART;TZID=America/Chicago:20251119T133000
DTEND;TZID=America/Chicago:20251119T135200
UID:submissions.supercomputing.org_SC25_sess302_pap841@linklings.com
SUMMARY:X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts 
 Architectures on HPC Platforms
DESCRIPTION:Yueming Yuan, Ahan Gupta, and Jianping Li (University of Illin
 ois Urbana-Champaign); Sajal Dash and Feiyi Wang (Oak Ridge National Labor
 atory (ORNL)); and Minjia Zhang (University of Illinois Urbana-Champaign)\
 n\nEmerging expert-specialized Mixture-of-Experts (MoE) architectures, suc
 h as DeepSeek-MoE, deliver strong model quality through fine-grained exper
 t segmentation and large top-k routing. However, their scalability is limi
 ted by substantial activation memory overhead and costly all-to-all commun
 ication. Furthermore, current MoE training systems—primarily optimized for
  NVIDIA GPUs—perform suboptimally on non-NVIDIA platforms, leaving signifi
 cant computational potential untapped.\n\nIn this work, we present X-MoE, 
 a novel MoE training system designed to deliver scalable training performa
 nce for next-generation MoE architectures. X-MoE achieves this via several
  novel techniques, including efficient padding-free MoE training with cros
 s-platform kernels, redundancy-bypassing dispatch, and hybrid parallelism 
 with sequence-sharded MoE blocks. Our evaluation on the Frontier supercomp
 uter, powered by AMD MI250X GPUs, shows that X-MoE scales DeepSeek-style M
 oEs up to 545 billion parameters across 1,024 GPUs—10x larger than the lar
 gest trainable model with existing methods under the same hardware budget,
  while maintaining high training throughput.\n\nTag: Best Student Paper Fi
 nalist, HPC for Machine Learning, Performance Measurement, Modeling, & Too
 ls\n\nRecording: Livestreamed, Recorded\n\nRegistration Category: Technica
 l Program Reg Pass\n\nSession Chair: Wahid Bhimji (Lawrence Berkeley Natio
 nal Laboratory (LBNL))\n\n
END:VEVENT
END:VCALENDAR
