Presentation
A Sample-Free Compilation Framework for Efficient Dynamic Tensor Computation
DescriptionDynamic-shape tensor computation poses challenges for shape-specific compilation due to variable input dimensions. Existing compilers rely on shape samples, incurring high tuning costs and degraded performance on unseen inputs.
We present Helix, a dynamic tensor framework with sample-free and architecture-guided compilation for compilation efficiency and shape-general performance. To avoid shape sampling, Helix constructs shape-agnostic compilation by decomposing computations across architectural layers. A bidirectional strategy combines top-down abstraction, aligning tensor computations with architectural hierarchies, and bottom-up kernel construction, building efficient execution strategies from reusable, architecture-aligned micro-kernels. A hybrid analyzer ensures accuracy through profiling at lower architectural levels, and achieves scalability through architecture-informed modeling at higher levels and runtime.
This hierarchical design eliminates shape-specific tuning and enables shape-adaptive execution. Evaluations on x86 CPUs, ARM CPUs, and NVIDIA GPUs demonstrate that Helix reduces compilation time by 174x over existing compilers and delivers 2.26x and 3.29x speedups over vendor libraries and dynamic-shape compilers, respectively.
We present Helix, a dynamic tensor framework with sample-free and architecture-guided compilation for compilation efficiency and shape-general performance. To avoid shape sampling, Helix constructs shape-agnostic compilation by decomposing computations across architectural layers. A bidirectional strategy combines top-down abstraction, aligning tensor computations with architectural hierarchies, and bottom-up kernel construction, building efficient execution strategies from reusable, architecture-aligned micro-kernels. A hybrid analyzer ensures accuracy through profiling at lower architectural levels, and achieves scalability through architecture-informed modeling at higher levels and runtime.
This hierarchical design eliminates shape-specific tuning and enables shape-adaptive execution. Evaluations on x86 CPUs, ARM CPUs, and NVIDIA GPUs demonstrate that Helix reduces compilation time by 174x over existing compilers and delivers 2.26x and 3.29x speedups over vendor libraries and dynamic-shape compilers, respectively.
Event Type
Paper
TimeTuesday, 18 November 202511:15am - 11:37am CST
Location275
HPC for Machine Learning



