BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201805Z
LOCATION:123
DTSTART;TZID=America/Chicago:20251117T133000
DTEND;TZID=America/Chicago:20251117T170000
UID:submissions.supercomputing.org_SC25_sess267_tut139@linklings.com
SUMMARY:Distributed Deep Learning on GPU-Based Clusters
DESCRIPTION:Prajwal Singhania, Lannie Dalton Hough, and Cunyang Wei (Unive
 rsity of Maryland)\n\nDeep learning (DL) is rapidly becoming pervasive in 
 almost all areas of computer science, and is even being used to assist com
 putational science simulations and data analysis. A key behavior of these 
 deep neural networks (DNNs) is that they reliably scale, i.e., they contin
 uously improve in performance when the number of model parameters and amou
 nt of data grow. As the demand for larger, more sophisticated, and more ac
 curate DL models increases, the need for large-scale parallel model traini
 ng, fine-tuning, and inference has become increasingly pressing. Subsequen
 tly, in the past few years, several parallel algorithms and frameworks hav
 e been developed to parallelize model training and inference on GPU-based 
 platforms. This tutorial will introduce and provide basics of the state of
  the art in distributed deep learning. We will use large language models (
 LLMs) as a running example, and teach the audience the fundamentals involv
 ed in performing the three essential steps of working with LLMs: (1) train
 ing an LLM from scratch, (2) continued training/fine-tuning of an LLM from
  a checkpoint, and (3) inference on a trained LLM. We will cover algorithm
 s and frameworks falling under the purview of data parallelism (PyTorch DD
 P and DeepSpeed), and tensor parallelism (AxoNN).\n\nRecording: Livestream
 ed, Recorded\n\nRegistration Category: Tutorial Reg Pass\n\n
END:VEVENT
END:VCALENDAR
