BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201809Z
LOCATION:230
DTSTART;TZID=America/Chicago:20251117T103000
DTEND;TZID=America/Chicago:20251117T110000
UID:submissions.supercomputing.org_SC25_sess202_ws_pdsw115@linklings.com
SUMMARY:LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing
  of Large Language Models
DESCRIPTION:Minqiu Sun (University of Delaware), Xin Huang (RIKEN Center f
 or Computational Science (R-CCS)), Luanzheng Guo and Nathan R. Tallent (Pa
 cific Northwest National Laboratory (PNNL)), Kento Sato (RIKEN Center for 
 Computational Science (R-CCS)), and Dong Dai (University of Delaware)\n\nC
 heckpointing is essential for fault tolerance in training large language m
 odels (LLMs). However, existing methods, regardless of I/O strategies, per
 iodically store the entire model and optimizer states, incurring substanti
 al storage overhead and contention. Recent studies reveal that updates acr
 oss LLM layers are highly non-uniform. During training, some layers may un
 dergo more significant changes, while others remain stable or even unchang
 ed. This suggests that selectively checkpointing only layers with signific
 ant updates could reduce overhead without harming training. Implementing s
 uch strategies requires fine-grained control over both weights and optimiz
 er states, which no current tool provides. To address this gap, we propose
  LLMTailor, a checkpoint-merging-framework that filters and assembles laye
 rs from different checkpoints to form a composite checkpoint. Our evaluati
 on indicates that LLMTailor can work with different  checkpointing strateg
 ies and effectively reduce checkpoint size (e.g., 4.3 times smaller for Ll
 ama3.1-8B) and checkpoint time (e.g., 2.8 times faster for Qwen2.5-7B) whi
 le maintaining model quality.\n\nTag: Data Analytics, High Performance I/O
 , Storage, Archive, & File Systems, Storage\n\nRecording: Livestreamed, Re
 corded\n\nRegistration Category: Technical Program Reg Pass, Workshop Reg 
 Pass\n\nSession Chairs: Suren Byna (The Ohio State University, Lawrence Be
 rkeley National Laboratory (LBNL)); Anthony Kougkas (Illinois Institute of
  Technology, Argonne National Laboratory (ANL)); Sarah Neuwirth (Johannes 
 Gutenberg University Mainz); Ricardo Macedo (INESC TEC, University of Minh
 o); Jay Lofstead (Sandia National Laboratories, University of New Mexico);
  and Dean Hildebrand (Google LLC)\n\n
END:VEVENT
END:VCALENDAR
