BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201300Z
LOCATION:Second Floor Atrium
DTSTART;TZID=America/Chicago:20251121T080000
DTEND;TZID=America/Chicago:20251121T120000
UID:submissions.supercomputing.org_SC25_sess620_post262@linklings.com
SUMMARY:Accelerating Linear Solve with Mixed Precision Nested Recursive Su
 bdivision on AI Hardware
DESCRIPTION:Vicki Carrica (Massachusetts Institute of Technology (MIT))\n\
 nThe Cholesky decomposition is a critical performance bottleneck in engine
 ering simulations. To accelerate these simulations, we present a novel, ne
 sted recursive Cholesky algorithm implemented in Julia. The algorithm rest
 ructures the problem into recursive TRSM (triangular solve) and SYRK (symm
 etric rank-k update) sub-problems, maximizing the use of highly parallel G
 EMM (general matrix-matrix multiply) operations that are highly efficient 
 on GPUs. This approach leverages a custom recursive data structure that en
 ables layered, mixed-precision arithmetic on modern NVIDIA H200 GPUs. By s
 trategically using fast, low-precision FP16 computations on large, off-dia
 gonal matrix blocks via Tensor Cores, while preserving high-precision on t
 he critical diagonal blocks, we achieve a speedup of 5.32x over the standa
 rd cuSOLVER FP64 implementation. This method is 100x more accurate than a 
 pure FP16 approach while retaining over 88\% of its speedup. Our work demo
 nstrates a practical path to significantly reducing computation time for l
 arge-scale scientific problems with minimal accuracy loss.\n\nTag: Researc
 h & ACM SRC Posters\n\nRegistration Category: Technical Program Reg Pass\n
 \nSession Chairs: Kento Sato (RIKEN Center for Computational Science (R-CC
 S)); Anja Gerbes (Georg-August-Universität Göttingen); and Chris Schlipali
 us (Pawsey Supercomputing Research Centre; Commonwealth Scientific and Ind
 ustrial Research Organisation (CSIRO), Australia)\n\n
END:VEVENT
END:VCALENDAR
