BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201803Z
LOCATION:265
DTSTART;TZID=America/Chicago:20251117T155500
DTEND;TZID=America/Chicago:20251117T162000
UID:submissions.supercomputing.org_SC25_sess217_ws_drbsd113@linklings.com
SUMMARY:Compression Error Sensitivity Analysis for Different Experts in Mo
 E Model Inference
DESCRIPTION:Songkai Ma (Hong Kong Polytechnic University); Zhaorui Zhang (
 The Hong Kong Polytechnic University); Sheng Di (Argonne National Laborato
 ry (ANL)); Benben Liu (The University of Hong Kong); Xiaodong Yu (Stevens 
 Institute of Technology); Xiaoyi Lu (University of California, Merced); an
 d Dan Wang (Hong Kong Polytechnic University)\n\nWith the widespread appli
 cation of Mixture of Experts (MoE) reasoning models in the field of LLM le
 arning, efficiently serving MoE models under limited GPU memory constraint
 s has emerged as a significant challenge. Offloading the non-activated exp
 erts to main memory has been identified as an efficient approach to addres
 s such a problem, while it brings the challenges of transferring the exper
 t between the GPU memory and main memory. We need to explore an efficient 
 approach to compress the expert and analyze how the compression error affe
 cts the inference performance. \n\nTo bridge this gap, we propose employin
 g error-bounded lossy compression algorithms (such as SZ3 and CuSZp) to co
 mpress non-activated experts, thereby reducing data transfer overhead duri
 ng MoE inference. We conduct extensive experiments across various benchmar
 ks and present a comprehensive analysis of how compression-induced errors 
 in different experts affect overall inference accuracy.\n\nRecording: Live
 streamed, Recorded\n\nRegistration Category: Technical Program Reg Pass, W
 orkshop Reg Pass\n\nSession Chairs: Sheng Di (Argonne National Laboratory 
 (ANL), University of Chicago); Ana Gainaru (Oak Ridge National Laboratory 
 (ORNL)); Kento Sato (RIKEN Center for Computational Science (R-CCS)); Xin 
 Liang (University of Kentucky); and Jieyang Chen (University of Oregon)\n\
 n
END:VEVENT
END:VCALENDAR
