BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201229Z
LOCATION:Second Floor Atrium
DTSTART;TZID=America/Chicago:20251118T080000
DTEND;TZID=America/Chicago:20251118T170000
UID:submissions.supercomputing.org_SC25_sess537_drs115@linklings.com
SUMMARY:Improving Collective Aggregation for HPC and AI Workloads
DESCRIPTION:Mikaila Gossman (Clemson University)\n\nHigh performance compu
 ting (HPC) applications generate massive volumes of data, placing sustaine
 d pressure on parallel file systems (PFS) that face limited bandwidth and 
 resource contention. While file-per-process I/O allows lock-free access, r
 educing stripe contention, it creates excessive metadata overhead and poor
  manageability at scale. Aggregation—consolidating output from many proces
 ses into fewer shared files—helps mitigate these issues, but introduces ne
 w challenges related to concurrency, resource contention, complex I/O patt
 erns, and their interactions with heterogeneous storage devices.\n\nWe ide
 ntify and evaluate key I/O bottlenecks across these dimensions. To support
  system-level tuning, we introduce a lightweight OpenMP benchmark that hel
 ps users identify optimal aggregation parameters and found that interleave
 d, append-only I/O provides better performance when aggregating to a share
 d file. From this work, we present a novel, producer-consumer-based aggreg
 ation model designed to balance concurrency and resource usage efficiently
 . In microbenchmarks, our strategy achieved up to 2× higher write throughp
 ut than GIO and 1.6× higher than ADIOS2. In a real-world HPC application (
 HACC), it delivered 1.2× higher throughput with only 3% checkpoint overhea
 d—compared to ~12% for GIO, which is optimized for HACC. Finally, we demon
 strate the limitations of existing checkpointing approaches using DeepSpee
 d Megatron on the BLOOM 3B model, revealing significant inefficiencies dur
 ing restore phase due to excessive reads and seeks.\n\nFuture work will ex
 tend our aggregation framework for large language model (LLM) C/R, which i
 ntroduces highly concurrent, small, and random I/O patterns that pose new 
 challenges for traditional PFS architectures.\n\nTag: Research & ACM SRC P
 osters\n\nRecording: Not Livestreamed, Not Recorded\n\nRegistration Catego
 ry: Technical Program Reg Pass\n\nSession Chairs: Kento Sato (RIKEN Center
  for Computational Science (R-CCS)); Chris Schlipalius (Pawsey Supercomput
 ing Research Centre; Commonwealth Scientific and Industrial Research Organ
 isation (CSIRO), Australia); and Anja Gerbes (Georg-August-Universität Göt
 tingen)\n\n
END:VEVENT
END:VCALENDAR
