BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201805Z
LOCATION:275
DTSTART;TZID=America/Chicago:20251118T111500
DTEND;TZID=America/Chicago:20251118T113700
UID:submissions.supercomputing.org_SC25_sess290_pap169@linklings.com
SUMMARY:A Sample-Free Compilation Framework for Efficient Dynamic Tensor C
 omputation
DESCRIPTION:Yangjie Zhou (Tencent, National University of Singapore); Hong
 lin Zhu and Qian Qiu (Tencent); Weihao Cui (Shanghai Jiao Tong University)
 ; Zihan Liu (Shanghai Jiao Tong University, Shanghai Qi Zhi Institute); Pe
 ng Chen and Mohamed Wahib (RIKEN Center for Computational Science (R-CCS))
 ; Cong Guo and Siyuan Feng (Shanghai Jiao Tong University); Jintao Meng (S
 henzhen Institute of Advanced Technology, Chinese Academy of Sciences); Ha
 idong Lan (Taichi Graphics); Jingwen Leng (Shanghai Jiao Tong University, 
 Shanghai Qi Zhi Institute); Yun Lin (Shanghai Jiao Tong University); Jin S
 ong Dong (National University of Singapore); and Wenxi Zhu and Minwen Deng
  (Tencent)\n\nDynamic-shape tensor computation poses challenges for shape-
 specific compilation due to variable input dimensions. Existing compilers 
 rely on shape samples, incurring high tuning costs and degraded performanc
 e on unseen inputs.\n\nWe present Helix, a dynamic tensor framework with s
 ample-free and architecture-guided compilation for compilation efficiency 
 and shape-general performance. To avoid shape sampling, Helix constructs s
 hape-agnostic compilation by decomposing computations across architectural
  layers. A bidirectional strategy combines top-down abstraction, aligning 
 tensor computations with architectural hierarchies, and bottom-up kernel c
 onstruction, building efficient execution strategies from reusable, archit
 ecture-aligned micro-kernels. A hybrid analyzer ensures accuracy through p
 rofiling at lower architectural levels, and achieves scalability through a
 rchitecture-informed modeling at higher levels and runtime.\n\nThis hierar
 chical design eliminates shape-specific tuning and enables shape-adaptive 
 execution. Evaluations on x86 CPUs, ARM CPUs, and NVIDIA GPUs demonstrate 
 that Helix reduces compilation time by 174x over existing compilers and de
 livers 2.26x and 3.29x speedups over vendor libraries and dynamic-shape co
 mpilers, respectively.\n\nTag: HPC for Machine Learning, Programming Frame
 works\n\nRecording: Livestreamed, Recorded\n\nRegistration Category: Techn
 ical Program Reg Pass\n\nSession Chair: Lena Oden (Juelich Supercomputing 
 Centre, Fernuniversitaet Hagen)\n\n
END:VEVENT
END:VCALENDAR
