BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201229Z
LOCATION:Second Floor Atrium
DTSTART;TZID=America/Chicago:20251118T080000
DTEND;TZID=America/Chicago:20251118T170000
UID:submissions.supercomputing.org_SC25_sess537_drs113@linklings.com
SUMMARY:Exploring Efficient Deep Learning Training on AI Accelerators
DESCRIPTION:Milan Shah (North Carolina State University)\n\nThe computatio
 nal and memory demands of DNN training have grown with the size of AI mode
 ls in recent years. To address these demands, popular accelerators (i.e., 
 GPUs) must find novel ways to reduce memory utilization since their memory
  capacity is on the scale of tens of GB. Other companies have unveiled nov
 el AI accelerators, generally with high on-chip memory capacity and varyin
 g architectures. For these accelerators, frequent on-chip/off-chip memory 
 transactions can bottleneck performance. Lossy compression is a promising 
 tool to reduce data footprint for efficient DNN training. Our work studies
  lossy compressors targeting training data and activation data, and how to
  efficiently run compression and GNN training on novel AI accelerators.\n\
 nOur contributions are: 1) a novel, portable training data compressor, cal
 led DCT+Chop, for emerging AI accelerators; 2) an activation compression f
 ramework tailored to the Graphcore Intelligence Processing Unit (IPU); 3) 
 a GPU-based design for a compressor/optimizer-agnostic lossy activation co
 mpression framework, called LAT-ACT; and 4) an exploration in training gra
 ph neural networks (GNNs) on the Cerebras CS-2. DCT+Chop and IPU activatio
 n compression have yielded strong results, where DCT+Chop can compress tra
 ining data up to 16X with a throughput on the scale of tens of GB/s. IPU a
 ctivation compression can speedup single IPU training up to 3.5X and multi
 -IPU training by several orders of magnitude. Preliminary results suggest 
 LAT-ACT yields compression ratios of 4-12X with limited accuracy degradati
 on. GNN training on the CS-2 can be implemented with PyTorch APIs, but fur
 ther exploration is needed for supporting sparse operators common to GNNs.
 \n\nTag: Research & ACM SRC Posters\n\nRecording: Not Livestreamed, Not Re
 corded\n\nRegistration Category: Technical Program Reg Pass\n\nSession Cha
 irs: Kento Sato (RIKEN Center for Computational Science (R-CCS)); Chris Sc
 hlipalius (Pawsey Supercomputing Research Centre; Commonwealth Scientific 
 and Industrial Research Organisation (CSIRO), Australia); and Anja Gerbes 
 (Georg-August-Universität Göttingen)\n\n
END:VEVENT
END:VCALENDAR
