BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201803Z
LOCATION:230
DTSTART;TZID=America/Chicago:20251117T160500
DTEND;TZID=America/Chicago:20251117T161000
UID:submissions.supercomputing.org_SC25_sess202_ws_pdswwip108@linklings.co
 m
SUMMARY:LLM training in practice: insights from 85,000 checkpoints
DESCRIPTION:Glenn K. Lockwood (VAST Data)\n\nTraining large language model
 s (LLMs) at scale generates significant I/O, and vendor guidance typically
  recommends provisioning performance based on the supply side: peak bandwi
 dth required to keep GPUs busy. These recommendations often overstate requ
 irements though, since they assume ideal GPU utilization. The demand side,
  or the I/O performance that training jobs actually drive, is not as well 
 characterized. Drawing on telemetry from production VAST systems underpinn
 ing some of the world’s largest AI training supercomputers, we analyzed ov
 er 85,000 checkpoints from 40 production LLM training jobs and found that 
 even trillion-parameter models require only a few hundred GB/s for efficie
 nt checkpointing. From these observations, we derive a simple, demand-side
  model that relates LLM size and checkpoint interval to the global bandwid
 th needed. This model offers a way to avoid overprovisioning I/O and to ma
 ximize the resources (power, cooling) that can go towards compute.\n\nTag:
  Data Analytics, High Performance I/O, Storage, Archive, & File Systems, S
 torage\n\nRecording: Livestreamed, Recorded\n\nRegistration Category: Tech
 nical Program Reg Pass, Workshop Reg Pass\n\nSession Chairs: Suren Byna (T
 he Ohio State University, Lawrence Berkeley National Laboratory (LBNL)); A
 nthony Kougkas (Illinois Institute of Technology, Argonne National Laborat
 ory (ANL)); Sarah Neuwirth (Johannes Gutenberg University Mainz); Ricardo 
 Macedo (INESC TEC, University of Minho); Jay Lofstead (Sandia National Lab
 oratories, University of New Mexico); and Dean Hildebrand (Google LLC)\n\n
END:VEVENT
END:VCALENDAR
