BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201258Z
LOCATION:Second Floor Atrium
DTSTART;TZID=America/Chicago:20251121T080000
DTEND;TZID=America/Chicago:20251121T120000
UID:submissions.supercomputing.org_SC25_sess620_post126@linklings.com
SUMMARY:Heterogeneity-Aware Task Allocation for Modern HPC Systems
DESCRIPTION:Sowmya Yellapragada (University of Utah); Jessica Imlau Dagost
 ini (University of California, Santa Cruz); and Kevin Gott and Rebecca Har
 tman-Baker (Lawrence Berkeley National Laboratory (LBNL))\n\nModern superc
 omputing systems exhibit heterogeneous node configurations, where seemingl
 y identical hardware exhibits significant performance variations due to me
 mory capacity differences, manufacturing tolerances, and deployment condit
 ions. This heterogeneity impacts the efficiency of scientific applications
  built on frameworks like AMReX, leading to substantial computational wast
 e on leadership-class systems. We present performance-aware and relation-a
 ware load balancing algorithms specifically designed for scientific applic
 ations, like AMReX on heterogeneous HPC clusters. Our approach uses empiri
 cally measured node performance characteristics and a relative performance
  matrix to optimize task distribution across diverse computational resourc
 es.\n\nEvaluation of NERSC Perlmutter with 14 representative AMReX computa
 tional kernels demonstrates 99.9% scheduling efficiency, achieving perform
 ance improvements of 4.4%-11.5% over traditional methods in moderate heter
 ogeneity scenarios (A100 40GB vs. 80GB) and up to 300x improvements in ext
 reme CPU-GPU mixed configurations where homogeneous methods fail to utiliz
 e CPU resources effectively. The algorithms handle million-task workloads 
 with O(nlogn + nm) complexity while maintaining practical deployment feasi
 bility.\n\nTag: Research & ACM SRC Posters\n\nRegistration Category: Techn
 ical Program Reg Pass\n\nSession Chairs: Kento Sato (RIKEN Center for Comp
 utational Science (R-CCS)); Anja Gerbes (Georg-August-Universität Göttinge
 n); and Chris Schlipalius (Pawsey Supercomputing Research Centre; Commonwe
 alth Scientific and Industrial Research Organisation (CSIRO), Australia)\n
 \n
END:VEVENT
END:VCALENDAR
