Presentation
High-Performance and Smart Networking Technologies for HPC and AI
DescriptionHigh-performance networking technologies are generating a lot of excitement towards building next-generation high-end computing (HEC) systems for HPC and AI with GPUs, accelerators, data center processing units (DPUs), and a variety of application workloads. This tutorial provides an overview of these continuously evolving technologies, their architectural features, current market standing, and suitability for designing HEC systems. We present a bottom-up view of various major scale-out interconnects (IB, HSE, RoCE, Omni-Path, AWS-EFA, Cray/HPE Slingshot, and Fujitsu Tofu-D) as well as scale-up interconnects like NVLink/NVSwitch and AMD Infinity Fabric. Integration of these technologies into libraries such as UCX and Libfabric is also discussed. Emerging standards like Ultra Ethernet, UALink, and Scale-Up Ethernet (SUE) are also presented. Next, we provide an overview of GPU Direct RDMA technology, DPU/IPU technology (NVIDIA BlueField, AMD Pensando, Intel IPUs), and AI-specific hardware (Cerebras Wafer-Scale Engines and Intel/Habana-Gaudi processors). Finally, we provide an overview of sample performance numbers that can be harnessed from these networking technologies. The tutorial also includes a set of hands-on exercises to help attendees understand these technologies from the ground up, following the flow of the tutorial (networking technologies, MPI library integration, GPU-Awareness in MPI libraries, and DPU-Awareness in MPI libraries).
Event Type
Tutorial
TimeMonday, 17 November 20258:30am - 5:00pm CST
Location130
Livestreamed
Recorded





