BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201809Z
LOCATION:131
DTSTART;TZID=America/Chicago:20251116T083000
DTEND;TZID=America/Chicago:20251116T170000
UID:submissions.supercomputing.org_SC25_sess269_tut150@linklings.com
SUMMARY:Principles and Practice of High-Performance Deep/Machine Learning 
 Training and Inference
DESCRIPTION:Dhabaleswar K. (DK) Panda, Hari Subramoni, and Nawras Alnaasan
  (The Ohio State University) and Jinghan Yao, Chen Chun Chen, and Lang Xu 
 (Ohio State University)\n\nRecent advances in machine learning and deep le
 arning (ML/DL) have led to many exciting challenges and opportunities. Mod
 ern ML/DL frameworks including PyTorch, TensorFlow, and cuML enable high-p
 erformance training, inference, and deployment for various types of ML mod
 els and deep neural networks (DNNs). This tutorial provides an overview of
  recent trends in ML/DL and the role of cutting-edge hardware architecture
 s and interconnects in moving the field forward. We will also present an o
 verview of different DNN architectures, ML/DL frameworks, DL training and 
 inference, and hyperparameter optimization, with special focus on parallel
 ization strategies for large models such as GPT, LLaMA, DeepSeek, and ViT.
  We highlight new challenges and opportunities for communication runtimes 
 to exploit high-performance CPU/GPU architectures to efficiently support l
 arge-scale distributed training. We also highlight some of our co-design e
 fforts to utilize MPI for large-scale DNN training on cutting-edge CPU/GPU
 /DPU architectures available on modern HPC clusters. Throughout the tutori
 al, we include several hands-on exercises to enable attendees to gain firs
 thand experience of running distributed ML/DL training and hyperparameter 
 optimizations on a modern GPU cluster.\n\nRecording: Livestreamed, Recorde
 d\n\nRegistration Category: Tutorial Reg Pass\n\n
END:VEVENT
END:VCALENDAR
