BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201808Z
LOCATION:263-264
DTSTART;TZID=America/Chicago:20251119T133000
DTEND;TZID=America/Chicago:20251119T135200
UID:submissions.supercomputing.org_SC25_sess289_pap809@linklings.com
SUMMARY:Fine-Grained Automated Failure Management for Extreme-Scale GPU-Ac
 celerated Systems
DESCRIPTION:Yonatan Levitt, Richard Barella, Sam Zeltner, Tom Musta, Lance
  Cheney, Gustavo Espinosa, and Olivier Franza (Intel Corporation) and Bala
 zs Gerofi (Intel Corporation, RIKEN Center for Computational Science (R-CC
 S))\n\nAs high performance computing (HPC) systems scale in size, system-w
 ide hardware failure rates increase. Historical data from previous large-s
 cale HPC installations illustrate this trend, with the mean time between f
 ailures (MTBF) decreasing steadily over the past decade. Recent studies fr
 om artificial intelligence and machine learning (AI/ML) training extrapola
 te MTBF declining even further for future GPU-accelerated systems. As MTBF
  decreases, mean time to repair (MTTR) becomes more pronounced, highlighti
 ng the need for efficient recovery strategies.\n\nThis paper presents an a
 utomated failure management system that addresses this issue by minimizing
  MTTR through real-time decision-making based on failure statistics. Our k
 ey contributions include a centralized meta-database for event history ana
 lysis including correlated events, fine-grained multi-strike repair polici
 es, and an automated recovery framework. Deployed on the Aurora supercompu
 ter, the proposed system has reduced MTTR by up to 84X compared to manual 
 servicing, leading to significant cost savings and decreased system downti
 me.\n\nTag: Algorithms, HPC for Machine Learning, Performance Measurement,
  Modeling, & Tools, State of the Practice\n\nRecording: Livestreamed, Reco
 rded\n\nRegistration Category: Technical Program Reg Pass\n\nSession Chair
 : Guanpeng Li (University of Florida)\n\n
END:VEVENT
END:VCALENDAR
