BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201337Z
LOCATION:Second Floor Atrium
DTSTART;TZID=America/Chicago:20251120T080000
DTEND;TZID=America/Chicago:20251120T170000
UID:submissions.supercomputing.org_SC25_sess533@linklings.com
SUMMARY:Poster Presentations (Research, ACM SRC Grads/Undergrads)
DESCRIPTION:High-Performance Sparse Attention on Tensor Cores: Fused3S and
  Beyond\n\nSparse attention is a core building block in many leading neura
 l network models, from graph-structured learning to sparse sequence modeli
 ng. It can be decomposed into a sequence of three sparse matrix operations
  (3S): sampled dense-dense matrix multiplication (SDDMM), softmax normaliz
 ation, and spar...\n\n\nZitong Li (University of California, Irvine)\n----
 -----------------\nDistributed Modular Digital Twin Network for High-Perfo
 rmance and Reliable Data Centers\n\nHigh performance computing (HPC) workl
 oads are driving rack power densities beyond 100 kW, creating unprecedente
 d stress on data center cooling and power systems. Conventional CFD-based 
 digital twins provide high-fidelity design optimization but are too comput
 ationally intensive and rigid for operat...\n\n\nYan Chen, Xing Lu, Cary F
 aulkner, Alex Vlachokostas, Hanlong Wan, and Jeremy Lerond (Pacific Northw
 est National Laboratory (PNNL))\n---------------------\nOptimizing the GPU
  All-Reduce Using Multiple Processes Per GPU\n\nLarge inter-GPU all-reduce
  operations, prevalent throughout deep learning, are bottlenecked by commu
 nication costs. Emerging heterogeneous architectures are comprised of comp
 lex nodes, often containing four GPUs and dozens to hundreds of CPU cores 
 per node. Parallel applications are typically accele...\n\n\nMichael Adams
  and Amanda Bienz (University of New Mexico)\n---------------------\nFrom 
 Petabytes to Predictions: Harnessing Large-Scale NeuroBlu Mental Health Da
 ta and ML To Mitigate Medication Non-Adherence\n\nMedication non-adherence
  is a major public health issue, especially within the behavioral health d
 omain, with traditional measurement methods often being unreliable. This s
 tudy uses a machine learning approach to predict medication adherence in a
  large cohort of over 446,000 patients with major depr...\n\n\nAlyson Coll
 ins, Cathy Sandoval, Maya Seshan, Srishti Srivastava, and Josh McWilliams 
 (University of Southern Indiana)\n---------------------\nDivide, Conquer, 
 and Denoise: Hybrid Parallel Diffusion with Memory-Aware Coarse-to-Fine In
 ference\n\nDiffusion models create high-quality images but are slow becaus
 e denoising steps run in sequence. We present a hybrid parallel diffusion 
 framework speeding up generation on mixed-capacity GPUs while keeping imag
 es coherent. First, we split each image into patches sized by each GPU’s m
 emory (i....\n\n\nFarhana Amin (Virginia Tech), Kanchon Gharami (Embry-Rid
 dle Aeronautical University), and Dimitrios Nikolopoulos (Virginia Tech)\n
 ---------------------\nJob Grouping-Based Intelligent Resource Recommendat
 ion Framework\n\nIn the current large-scale computing systems, users from 
 various scientific backgrounds submit batch jobs with a set of requested r
 esources. Manual resource selection in HPC facilities leads to early job t
 erminations and out-of-memory errors due to underestimation of resources, 
 or compute and memory...\n\n\nBeste Oztop (Boston University); Benjamin Sc
 hwaller, Vitus J. Leung, and Jim Brandt (Sandia National Laboratories); an
 d Brian Kulis, Manuel Egele, and Ayse K. Coskun (Boston University)\n-----
 ----------------\nExplicit Low-Order Finite-Element Wave Simulation Accele
 rated with Variable-Precision Computing Using INT8 Tensor Cores\n\nUsing l
 ow-precision cores for acceleration of PDE-based simulations with sparse o
 r small matrices is often challenging due to the frequent data conversion 
 between high- and low-precision variables, and that the required precision
  varies in time/space due to the heterogeneity of the target problem. A...
 \n\n\nKohei Fujita and Tsuyoshi Ichimura (The University of Tokyo, RIKEN);
  Muneo Hori (Japan Agency for Marine-Earth Science and Technology); and La
 lith Maddegedara (The University of Tokyo)\n---------------------\nJulia w
 ith Intelligent Runtime for Heterogeneous Computing\n\nJulia, a high-perfo
 rmance, high-level language, harnesses dynamic typing and LLVM’s Just-in-T
 ime compiler to match the speed of C and Fortran in production. Meanwhile,
  IRIS serves as a heterogeneous runtime that discovers devices dynamically
  and schedules concurrent work on CPUs, GPUs, FPGAs, ...\n\n\nNarasinga Ra
 o Miniskar, Pedro Valero-Lara, William Godoy, Keita Teranishi, and Jeffrey
  S. Vetter (Oak Ridge National Laboratory (ORNL))\n---------------------\n
 Tensor Core Accelerated Fast Multipole Method for GROMACS\n\nThe evaluatio
 n of long-range pairwise electrostatic forces is the most computationally 
 intensive component of molecular dynamics (MD) simulations. The fast multi
 pole method (FMM) is an alternative to reduce the computational complexity
 . In this work, we implemented a hybrid parallel FMM with MPI and...\n\n\n
 Jiamian Huang (Institute of Science Tokyo), Muhammad Umair Sadiq (KTH Roya
 l Institute of Technology), Rio Yokota (Institute of Science Tokyo), and B
 erk Hess (KTH Royal Institute of Technology)\n---------------------\nMemor
 y-Efficient CFD Based on MPS: Effective One-Billion-Cell Resolution on a S
 ingle Node\n\nWe investigate matrix product states (MPS), a tensor-network
  compression method, as a memory-efficient representation of flow variable
 s. A three-dimensional incompressible Navier-Stokes solver is implemented 
 entirely in MPS form and is applied to canonical flow problems. Results sh
 ow substantial mem...\n\n\nJunya Onishi (RIKEN Center for Computational Sc
 ience (R-CCS)); Ayato Takii (Kobe University, Japan; RIKEN Center for Comp
 utational Science (R-CCS)); Sangwon Kim (RIKEN Center for Computational Sc
 ience (R-CCS)); Younghwa Cho (Hokkaido University, Japan); and Makoto Tsub
 okura (Kobe University, Japan; RIKEN Center for Computational Science (R-C
 CS))\n---------------------\nUnderstanding GPU Utilization Using LDMS Data
  on Perlmutter\n\nGPGPU-based clusters and supercomputers have grown signi
 ficantly in popularity over the past decade. While numerous GPGPU hardware
  counters are available to users, their potential for workload characteriz
 ation remains underexplored. In this work, we analyze previously overlooke
 d GPU hardware counter...\n\n\nOnur Cankur (University of Maryland), Brian
  Austin (Lawrence Berkeley National Laboratory (LBNL)), and Abhinav Bhatel
 e (University of Maryland)\n---------------------\nAn Efficient GEMM Accel
 eration Method for LLM Inference with Variable-Length Sequences\n\nTransfo
 rmer-based large language models (LLMs) have demonstrated remarkable capab
 ilities in natural language processing (NLP) tasks. The transformer layer 
 in LLM involves substantial general matrix multiplication (GEMM). However,
  the sequence length variability leads to redundant computation and har...
 \n\n\nYu Zhang and Lu Lu (South China University of Technology)\n---------
 ------------\nWiCAT: Reducing Congestion at Wireless Interfaces in Heterog
 eneous Architectures\n\nHeterogeneous architectures integrating CPUs, GPUs
 , and memory controllers generate diverse traffic patterns that stress the
  on-chip network. Wireless networks-on-chip (WNoCs) provide fast, single-h
 op communication across distant nodes. However, their effectiveness is lim
 ited by congestion at wirele...\n\n\nTarun Sharma (IIIT Delhi)\n----------
 -----------\nScalable Multi-Node Multi-GPU Datalog Engine with Energy-Awar
 e Profiling\n\nExascale computing, powered by GPUs, is reshaping high-perf
 ormance computing. Declarative languages such as Datalog naturally benefit
  from this shift, as recursive rules can be compiled into GPU-optimized re
 lational operations. Unlike SQL, Datalog executes queries iteratively unti
 l a fixed point is ...\n\n\nAhmedur Rahman Shovon (Argonne National Labora
 tory (ANL)) and Sidharth Kumar (University of Illinois Chicago)\n---------
 ------------\nAlgorithms and Applications of Dynamic Network Analysis Usin
 g CANDY\n\nMany complex systems across diverse domains can be represented 
 as dynamic networks, where entities are modeled as time-varying nodes and 
 interactions among these entities are modeled as evolving edges. Analyzing
  such networks provides insights into the underlying temporal characterist
 ics of the syst...\n\n\nAashish Pandey (University of North Texas), Arinda
 m Khanda and S.M. Shovan (Missouri University of Science and Technology), 
 Ali Y. Khan (University of North Texas), Boyana Norris (University of Oreg
 on), Sajal K. Das (Missouri University of Science and Technology), and San
 jukta Bhowmick (University of North Texas)\n---------------------\nTidalMa
 rk: A Scalable Benchmark for Coastal Water Level Forecasting\n\nAccurate f
 orecasting of water levels is essential for flood mitigation. Traditionall
 y, predictions have been based on harmonic analysis and sensor networks ma
 intained by the National Oceanographic and Atmospheric Administration. How
 ever, these methods struggle with high-variance events that change w...\n\
 n\nLucas Raicu, Daniel Grzenda, Ian Foster, and Kyle Chard (University of 
 Chicago)\n---------------------\nOptimizing and Extending Periodogram Comp
 utations for Astronomy\n\nThis work extends nifty-ls, a high-performance L
 omb–Scargle periodogram implementation, with multiple Fourier terms harmon
 ic fitting with OpenMP-parallelized/GPU-parallelized methods.\n\nWe levera
 ge a fast and accurate spreading kernel (the "exponential of semicircle") 
 from the Flatiron Institut...\n\n\nYuwei Sun (Flatiron Institute, Universi
 ty of Illinois Urbana-Champaign) and Lehman Garrison (Flatiron Institute)\
 n---------------------\nMitigating I/O Bottlenecks in LiDAR Pipelines by D
 irectly Merging Neural Decompression and Semantic Segmentation\n\nThe incr
 easing volume of high-resolution LiDAR data poses a significant I/O bottle
 neck in large-scale analysis and high-performance computing pipelines due 
 to costly intermediary data storage and retrieval. We introduce a novel, e
 nd-to-end framework that addresses this issue by proposing the first u...\
 n\n\nEthan Marquez, Max Faykus, Oyinlolu Odetoye, Melissa Smith, and Jon C
 alhoun (Clemson University)\n---------------------\nVaultX Merge: Breaking
  Memory Barriers in Proof-of-Space Plot Generation\n\nProof-of-work blockc
 hains, like Bitcoin, consume substantial energy, motivating greener altern
 atives such as proof-of-space (PoSp), which relies on storage rather than 
 computation. Existing PoSp implementations face scalability challenges due
  to high memory and I/O requirements, especially when gene...\n\n\nArnav S
 irigere, Varvara Bondarenko, and Ioan Raicu (Illinois Institute of Technol
 ogy)\n---------------------\ncsDF: A Double-Float Arithmetic Library for t
 he Cerebras CS-2\n\nRecently, there have been attempts to utilize AI accel
 erators for scientific computing; however, these devices generally lack ha
 rdware support for double-precision floating-point arithmetic, which is es
 sential for many scientific applications.\n\nThe Cerebras CS-2 system (CS-
 2) delivers extremely high...\n\n\nReo Nagashima, Akeru Nakamura, Kai Mura
 kami, and Ryunosuke Matsuzaki (Meiji University); Daichi Mukunoki (Nagoya 
 University); and Takaaki Miyajima (Meiji University)\n--------------------
 -\nEvaluating LiDAR Compression for 3D Semantic Segmentation in Diverse Of
 f-Road Environments on GOOSE Dataset\n\nTransmitting point cloud data is v
 ital for applications like autonomous vehicle navigation, especially for c
 ompute-limited vehicles. LiDAR data can easily grow to gigabytes or teraby
 tes uncompressed, making data transmission costly. While recent research h
 as advanced point cloud compression, most wo...\n\n\nAdam Niemczura, Max F
 aykus, Oyinlolu Odetoye, Melissa Smith, Jon Calhoun, and Scott Groel (Clem
 son University)\n---------------------\nMixed Compute Environments with Op
 enCHAMI\n\nThere is a growing need for workloads that don’t follow a tradi
 tional HPC workflow. Many of these workloads are developed with Kubernetes
  as the workload manager rather than an HPC-focused one such as Slurm. Mix
 ing different workloads presents a challenge for a few reasons: The demand
  for eith...\n\n\nSean Gibson, Richard Kim, Samuel Quan, Travis Cotton, an
 d Thomas Mackell (Los Alamos National Laboratory (LANL))\n----------------
 -----\nOptimizing Task-Driven Offloading in LLVM\n\nWe investigate an inef
 ficiency in the LLVM OpenMP runtime related to accelerator offloading. The
  current implementation manages asynchronous GPU tasks by polling async ha
 ndles, which introduces CPU overhead. We propose replacing this polling mo
 del with an event-driven approach that detaches target t...\n\n\nJan Kraus
 , Joachim Jenke, and Christian Terboven (Chair for High-Performance Comput
 ing i12, RWTH Aachen University)\n---------------------\nAn Agent-Based Vi
 ral Venture: Adaptive Tool Selection for Scalable Genomics\n\nFecal microb
 ial transplant (FMT) is an effective procedure for restoring gut microbiom
 e balance in patients with Clostridioides difficile infection by introduci
 ng healthy donor microbes. Tracking viral genomes during FMT provides insi
 ght into microbial community transfer and recovery. We developed a...\n\n\
 nNaomi Kolodisner (University of Arizona), Alok Kamatar (Advisor) (Univers
 ity of Chicago), and J. Greg Pauloski (Advisor) (NVIDIA Corporation)\n----
 -----------------\nPhySiViT: A Physics Simulation Vision Transformer\n\nMo
 dern scientific computing generates massive simulation data across physics
  domains, yet researchers lack general-purpose tools for efficient analysi
 s. While vision transformers like CLIP and DINO have revolutionized natura
 l image analysis, no equivalent exists for physics simulation data. This p
 ro...\n\n\nJessica Ezemba (Carnegie Mellon University), James Afful (Iowa 
 State University), and Mei-Yu Wang (Pittsburgh Supercomputing Center)\n---
 ------------------\nGATSched: Multi-Objective Graph Attention Networks for
  Energy-Efficient HPC Job Scheduling\n\nHigh performance computing (HPC) s
 ystems face an urgent sustainability crisis, with leading facilities consu
 ming 10–60 MW and incurring multimillion-dollar annual energy costs. Tradi
 tional schedulers like SLURM and PBS treat energy as secondary, leading to
  30%–50% energy waste above theo...\n\n\nKyrian Adimora (The University of
  Kansas)\n---------------------\nConfiguring Large Language Models for Reg
 ional Ocean Model Development\n\nRecent work at NSF NCAR has developed Pyt
 hon packages and documentation for instantiating regional ocean models in 
 the Community Earth System Model, but how can we guide a community of user
 s through the subsequent tuning and development of purpose-built models? H
 ere, we leverage recent advances in n...\n\n\nAidan Janney (National Cente
 r for Atmospheric Research (NCAR), University of Colorado Boulder); Giovan
 ni Seijo-Ellis (University of Puerto Rico, Mayaguez; National Center for A
 tmospheric Research (NCAR)); and Dan Amrhein (National Center for Atmosphe
 ric Research (NCAR))\n---------------------\nNovel Graph Alignment Algorit
 hms for Identifying Non-Determinism in Large-Scale Simulations\n\nThe incr
 easing complexity of HPC simulations poses several challenges to their rep
 roducibility and reliability. One critical issue is the non-determinism (N
 D) induced by asynchronous MPI communication. Locating the sources of ND i
 n large codes is difficult. This problem can be addressed by comparing...\
 n\n\nDhroov Pandey (University of North Texas)\n---------------------\nPar
 allel Local Motif Counting on Large-Scale Dynamic Graphs\n\nGraph motifs—s
 mall subgraphs such as triangles and cliques—are key tools for comparing a
 nd aligning networks in domains ranging from biology to social sciences. W
 hile recent advances enable motif counting in billion-edge networks, exist
 ing methods focus mainly on global frequencies. Buil...\n\n\nAli Khan and 
 Sanjukta Bhowmick (University of North Texas) and Michela Taufer (Universi
 ty of Tennessee, Knoxville)\n---------------------\nCharacterizing Perform
 ance and Energy Trade-Offs on the Aurora Supercomputer\n\nIn this work, we
  evaluate the range of performance-energy tradeoffs achievable by the inde
 pendent application of node-level power capping and DVFS controls on the A
 urora supercomputer. We then analyze the default uncore frequency behavior
  under constrained node power, revealing inefficiencies—...\n\n\nSolomon B
 ekele, Swann Perarnau, and Brice Videau (Argonne National Laboratory (ANL)
 )\n---------------------\nFacilitating Mixed Python-Fortran HPC Codes: 4D 
 Drift-Kinetic Simulations with Pyccel\n\nPython is widely used in scientif
 ic computing for prototyping, but its performance and memory overhead limi
 t its suitability for production in high-performance computing (HPC) envir
 onments. Pyccel addresses this by translating Python into human-readable F
 ortran or C, while retaining Python interoper...\n\n\nEmily Bourne (Swiss 
 Federal Institute of Technology Lausanne (EPFL), EPFL) and Yaman Güçlü (Ma
 x Planck Institute for Plasma Physics, Division of Numerical Methods in Pl
 asma Physics)\n---------------------\nA Toolbox for Load Balancing Develop
 ment and Analysis in WarpX/AMReX Applications\n\nEfficient load balancing 
 is critical for the scalability of distributed scientific applications. Ho
 wever, there are several challenges for applications to test new balancing
  strategies, including the need for an easy workflow to validate different
  algorithms. This work aims to tackle this particular...\n\n\nJessica Imla
 u Dagostini (University of California, Santa Cruz); Sowmya Yellapragada (U
 niversity of Utah); and Kevin Gott and Rebecca Hartman-Baker (Lawrence Ber
 keley National Laboratory (LBNL))\n---------------------\nScalable Alterna
 tive Route Computation with ACE: A C++17 Library for HPC Traffic Simulatio
 ns\n\nWe present ACE (Asynchronous Communication and Execution), a C++17 l
 ibrary for scalable asynchronous task execution on high performance comput
 ing (HPC) systems. Integrated into a distributed traffic simulation workfl
 ow, ACE accelerates the computation of alternative routes, a key performan
 ce bottlen...\n\n\nPaulo Silva, Pavlína Smolková, Kateřina Slaninová, Jan 
 Martinovič, João Barbosa, and Matej Špeťko (IT4Innovations, VSB - Technica
 l University of Ostrava) and Emanuele Vitali (CSC - IT Center for Science)
 \n---------------------\nSeamless Scaling of Applications Across Programmi
 ng Models\n\nWe present a comparative study of the productivity and perfor
 mance of four programming languages: Python, Julia, C++, and DaphneDSL, fo
 r the Connected Components graph algorithm from the GAP benchmark suite. U
 sing various code productivity metrics, we evaluated the effort of scaling
  applications fro...\n\n\nReto Krummenacher (University of Basel), Quentin
  Guilloteau (Inria), Jonas H. Müller Korndörfer (University of Bern), and 
 Florina M. Ciorba (University of Basel)\n---------------------\nProcess-Ba
 sed Predictors of Vulnerability Reintroduction\n\nThere is growing interes
 t in securing scientific software, which underpins research results and of
 ten transitions into commercial systems. While source code metrics provide
  useful indicators of vulnerabilities, software engineering process (SEP) 
 metrics can uncover patterns that lead to their introd...\n\n\nSamiha Shim
 mi (Northern Illinois University), Nicholas Synovic (Loyola University Chi
 cago), Mona Rahimi (Northern Illinois University), and George Thiruvathuka
 l (Loyola University Chicago)\n---------------------\nEnhancing Usability 
 and Performance in Experimental Environments Management\n\nReproducibility
  is a challenge in HPC and research. HPC experiments are resource-intensiv
 e and depend on complex software environments. Snapshotting addresses this
  issue, captures the complete state of a system in a single step, allowing
  researchers to automatically rebuild and restore identical env...\n\n\nZa
 hra Temori (University of Delaware), Paul Marshall (UChicago Department of
  Computer Science), and Kate Keahey (Argonne National Laboratory (ANL))\n-
 --------------------\nCIRE: LLVM Analysis for Floating-Point Rounding Erro
 r Affected by Precision and Optimizations\n\nNumerical programmers often a
 djust precision settings and compiler optimizations to maximize performanc
 e, but these changes can unpredictably affect floating-point rounding erro
 rs. We present CIRE, a tool that statically estimates tight bounds on floa
 ting-point rounding error by analyzing LLVM code ...\n\n\nCayden Lund, Tan
 may Tirpankar, and Ganesh Gopalakrishnan (University of Utah)\n-----------
 ----------\nEchoes of Earth: Building an Autonomous Environmental Lab for 
 Acoustic Sensing\n\nCurrent bioacoustic monitoring technologies cost $600-
 $1,000+ per device and require manual data retrieval and maintenance by ex
 perts, preventing real-time insights and limiting deployment scale. We dev
 elop a prototype autonomous monitoring and detection system that streams h
 igh-quality audio in rea...\n\n\nHudson Reynolds (Boston University); Alex
  Tuecke (Worcester Polytechnic Institute); Mike Sherman (University of Chi
 cago); and Kate Keahey (Argonne National Laboratory (ANL), University of C
 hicago)\n---------------------\nA Kokkos-Based Proxy of the Exascale Metag
 enome Assembler MetaHipMer2: A First Use of Kokkos for Computational Biolo
 gy\n\nInexpensive DNA sequencing [1] has opened new windows into biologica
 l complexity. These include metagenomics: the ability to catalog a microbi
 al ecosystem by extracting and sequencing DNA directly from an environment
 . Analyzing metagenomic-scale datasets often requires exascale computing. 
 Such compu...\n\n\nLogan Williams, Gavin Conant, and Michela Becchi (North
  Carolina State University) and Jan Ciesko and Amy Powell (Sandia National
  Laboratories)\n---------------------\nGNNs on Evolving Graphs: A Benchmar
 k of Incremental Updates and Meta-Learning Approaches\n\nThe field of grap
 h machine learning has seen significant growth with the success of graph n
 eural networks (GCNs). However, most traditional GCNs are designed for sta
 tic graphs. In the real world, graphs are constantly evolving—new users jo
 in social networks, molecules change shape, and data st...\n\n\nSriram Sri
 nivasan (Bowie State University), Sanjukta Bhowmick (University of North T
 exas), and Hamdan Alabsi and Rand Obeidat (Bowie State University)\n------
 ---------------\nC++ Standard Parallelism for GPU Programming in a Particl
 e-In-Cell Application\n\nPerformance portability remains a major challenge
  in high performance computing as applications increasingly target diverse
  GPU architectures. The C++17 standard introduced stdpar, a high-level par
 allelism model to simplify parallel programming. NVIDIA extended this mode
 l for GPU execution within he...\n\n\nEster El Khoury, Mathieu Lobet, and 
 Julien Bigot (CEA Saclay) and Laurent Colombet (CEA Dam)\n----------------
 -----\nFast Linear Solvers via AI-Tuned Markov Chain Monte Carlo-Based Mat
 rix Inversion\n\nLarge, sparse linear systems are pervasive in modern scie
 nce and engineering, and Krylov subspace solvers are an established means 
 of solving them. Yet convergence can be slow for ill-conditioned matrices,
  so practical deployments usually require preconditioners. Markov chain Mo
 nte Carlo (MCMC)-base...\n\n\nAnton Lebedev and Won Kyung Lee (STFC Hartre
 e Centre); Soumyadip Ghosh (IBM Thomas J. Watson Research Center); Olha I.
  Yaman (STFC Hartree Centre); Vassilis Kalantzis, Yingdong Lu, Tomasz Nowi
 cki, Shashanka Ubaru, and Lior Horesh (IBM Thomas J. Watson Research Cente
 r); and Vassil Alexandrov (STFC Hartree Centre)\n---------------------\nBe
 tween the NIC and a Hard Place: Evaluating 400 Gb/s Ethernet for HPC Data 
 Transfers\n\nThe readiness of new 400Gbps Ethernet hardware was evaluated 
 for potential production use in high performance computing (HPC) environme
 nts over a local area network (LAN) and a wide area network (WAN). The app
 roach explored a range of data movement strategies, including parallelized
  transfer tools, ...\n\n\nAdelle Ferris, Evelyn Needham, Nikole Grandez, J
 esse Martinez, and Doug Egan (Los Alamos National Laboratory (LANL))\n----
 -----------------\nGPU Kernels for Mixture of Experts\n\nThe Sparsely-Gate
 d Mixture of Experts (MoE) has seen a surge in use over the last year. Thi
 s is primarily motivated by a desire to increase the size of language mode
 ls, without a proportional increase in the total number of FLOPs. Due to t
 heir popularity, there is a large volume of work studying dis...\n\n\nArth
 ur Feeney (University of California, Irvine); Ying Wai Li (Los Alamos Nati
 onal Laboratory (LANL)); and Aparna Chandramowlishwaran (University of Cal
 ifornia, Irvine)\n---------------------\nSRAP: Sender-Side Receiver-Aware 
 Port Selection for High-Speed Multi-Flow TCP\n\nAchieving 400 Gbps require
 s aggregating multiple flows across cores rather than pushing single-flow 
 limits. However, with 16×25 Gbps TCP flows, random ephemeral ports cause r
 eceiver-side packet steering (RSS) to concentrate flows on single CPU core
 s, degrading throughput from 25 Gbps to below 5 Gbps...\n\n\nShingo Hattor
 i and Osamu Tatebe (University of Tsukuba)\n---------------------\nHardwar
 e-Aware Quantum Circuit Synthesis\n\nEffectively leveraging quantum comput
 ing requires generating and manipulating a desired quantum state using a q
 uantum circuit. Quantum circuit synthesis (QCS) is bottlenecked by the exp
 onential complexity of circuit verification via quantum simulation. Diffus
 ion models are promising QCS candidates, ...\n\n\nNathan Jones, Akhilesh B
 ondapalli, Toby Cox, Ian Lewis, and Rong Ge (Clemson University)\n--------
 -------------\nLuthier: A Dynamic Binary Instrumentation Framework Targeti
 ng AMD GPUs\n\nIn this poster we present Luthier, the first open-source dy
 namic binary instrumentation framework targeting AMD GPUs. We highlight ke
 y features of our framework, including example use cases and runtime overh
 ead comparison with NVIDIA’s NVBit. We also go over some major enhancement
 s under devel...\n\n\nMatin Raayai-Ardakani, Norman Rubin, and David Kaeli
  (Northeastern University)\n---------------------\nTemplate Task-Based Mul
 tiresolution Analysis in Hybrid Environments\n\nWe present a framework tha
 t implements multiresolution analysis (MRA) on top of Template Task Graph 
 (TTG), a distributed, task-based data-flow programming model. MRA is broad
 ly applied across scientific domains for its ability to capture both local
  and global features with high accuracy, and its ada...\n\n\nNilesh Chatur
 vedi (Institute for Advanced Computational Science, Stony Brook University
 ; Stony Brook University, Department of Applied Mathematics and Statistics
 ); Joseph Schuchart (Institute for Advanced Computational Science, Stony B
 rook University); and Robert J. Harrison (Institute for Advanced Computati
 onal Science, Stony Brook University; Stony Brook University, Department o
 f Applied Mathematics and Statistics)\n---------------------\nHeterogeneit
 y-Aware Task Allocation for Modern HPC Systems\n\nModern supercomputing sy
 stems exhibit heterogeneous node configurations, where seemingly identical
  hardware exhibits significant performance variations due to memory capaci
 ty differences, manufacturing tolerances, and deployment conditions. This 
 heterogeneity impacts the efficiency of scientific app...\n\n\nSowmya Yell
 apragada (University of Utah); Jessica Imlau Dagostini (University of Cali
 fornia, Santa Cruz); and Kevin Gott and Rebecca Hartman-Baker (Lawrence Be
 rkeley National Laboratory (LBNL))\n---------------------\nTowards a GPU-A
 ccelerated Web-Based Graph Rendering Framework for Large-Scale Protein Net
 works\n\nWe present a WebGPU-based framework for real-time visualization o
 f large-scale protein–protein interaction (PPI) networks directly in stand
 ard browsers. Built on GraphWaGu, our extended graph rendering API integra
 tes GPU-accelerated force-directed layout computation with dynamic edge fi
 ltering...\n\n\nJiaxin Lu and Landon Dyken (University of Illinois Chicago
 ); Shilpika Shilpika and Venkatram Vishwanath (Argonne National Laboratory
  (ANL)); Michael Papka (University of Illinois Chicago, Argonne National L
 aboratory (ANL)); and Sidharth Kumar (University of Illinois Chicago)\n---
 ------------------\nAccelerating Scientific Workflows with LLM-Driven Comp
 iler Optimizations for Generated High-Performance Hardware\n\nThe optimiza
 tion of computation kernels is central to high performance computing, dire
 ctly impacting applications from scientific computing to artificial intell
 igence (AI). In experimental workflows with high-throughput or streaming d
 ata, software-only execution often becomes a bottleneck, motivatin...\n\n\
 nRobert Ramstad, Nicolas Bohm Agostini, and Antonino Tumeo (Pacific Northw
 est National Laboratory (PNNL))\n---------------------\nA Formal Character
 ization of Non-Monotonicity in Tensor Cores\n\nModern high performance com
 puting increasingly relies on hardware accelerators like NVIDIA Tensor Cor
 es, which employ non-standard internal arithmetic that can evolve between 
 hardware generations. This non-standard approach can violate the fundament
 al mathematical property of monotonicity, leading t...\n\n\nPaul Jiang (Pu
 rdue University) and Vivian Zheng (Stony Brook University)\n--------------
 -------\nScODA: An Emerging Pipeline for Evaluating Distributed Database P
 erformance To Support Operational Data Analytics\n\nAs high performance co
 mputing (HPC) systems scale toward the exascale era, operational data anal
 ytics (ODA) play an increasingly central role in managing system security,
  health, scheduling, and scientific productivity. Supercomputing facilitie
 s continuously generate massive volumes of logs and syst...\n\n\nNicholas 
 Synovic (Loyola University Chicago); FNU Shilpika, Silvio Rizzi, and Doug 
 Waldron (Argonne National Laboratory (ANL)); George K. Thiruvathukal (Loyo
 la University Chicago); and Michael E. Papka (Argonne National Laboratory 
 (ANL))\n---------------------\nLocal vs. Global FFT Approaches for High-Pe
 rformance Ultrasound Simulation on Multi-GPU Systems\n\nSimulating wave pr
 opagation with the Fourier collocation method is computationally intensive
  due to its reliance on discrete Fourier transforms (DFTs). While DFTs ena
 ble near-minimal spatial discretization, they scale poorly on modern high 
 performance computing systems. This work evaluates two multi...\n\n\nOlive
 r Kuník and Jiri Jaros (Faculty of Information Technology, Brno University
  of Technology)\n---------------------\nAdversaGuard: A Distributed Data P
 oisoning Benchmark for Parallel AI\n\nThis study introduces FoodSAFE, a no
 vel high performance computing (HPC)-based distributed data poisoning (DDP
 ) framework designed to benchmark adversarial resilience and training perf
 ormance. The framework is tested across eight distinct configurations—seve
 n distributed frameworks and one non...\n\n\nYulia Kumar (Kean University,
  Rutgers University); Solomon Thomas, Dejaun Gayle, and J. Jenny Li (Kean 
 University); and Dov Kruger (Rutgers University)\n---------------------\nE
 nabling Real-Time, Extreme-Scale Bayesian Inference: FFT-Based GPU-Acceler
 ated Matrix-Vector Products for Block-Triangular Toeplitz Matrices\n\nAdjo
 int-based, matrix-free Newton-Krylov methods have long been the gold stand
 ard for solving high-dimensional, ill-posed inverse problems. These method
 s require a pair of forward and adjoint PDE solves per iteration, usually 
 making them intractable for real-time inference and prediction. We present
 ...\n\n\nSreeram Venkat and Omar Ghattas (The University of Texas at Austi
 n)\n---------------------\nUsing Hardware Metrics To Understand Performanc
 e of the RAJA Performance Suite Kernels in Different GPU Modes on MI300A\n
 \nModern GPUs play a crucial role in accelerating a wide range of computat
 ional workloads. However, their performance is often limited by the memory
  access patterns of the kernels they execute. AMD’s MI300A APU supports mu
 ltiple logical GPU partitioning modes to optimize compute resource allocat
 ...\n\n\nAmr Abouelmagd (Tennessee Tech University) and Stephanie Brink, M
 ichael McKinsey, David Boehme, Jason Burmark, Brian Ryujin, Tom Scogland, 
 and Olga Pearce (Lawrence Livermore National Laboratory (LLNL))\n---------
 ------------\nFrom Legacy to Portable: An Agentic AI Workflow for Fortran 
 Code Translation and Cross-Architecture Optimization\n\nLegacy Fortran cod
 es remain central to many scientific applications but are poorly suited to
  today’s GPU-accelerated heterogeneous architectures. Manual porting to pe
 rformance-portable frameworks like Kokkos is time-consuming and requires d
 eep domain expertise, creating a major barrier to mode...\n\n\nSparsh Gupt
 a (Los Alamos National Laboratory (LANL), Franklin W. Olin College of Engi
 neering) and Kamalavasan Kamalakkannan, Maxim Moraru, Galen Shipman, and P
 atrick Diehl (Los Alamos National Laboratory (LANL))\n--------------------
 -\nCompute System Simulator: Modeling the Impact of Allocation Policy and 
 Hardware Reliability on HPC Cloud Resource Utilization\n\nWe have develope
 d a comprehensive simulation tool to model the launching, progression, and
  completion of virtual machines and corresponding workloads within a cloud
  cluster of arbitrary size. The simulator employs various policies to allo
 cate computational resources for these virtual machines, simul...\n\n\nJar
 rod Leddy and Huseyin Yildiz (Microsoft Corporation)\n--------------------
 -\nApplying Lossy Compression Techniques to GNN Training\n\nGraph neural n
 etworks (GNNs) are a state-of-the-art machine learning model for processin
 g graph-structured data. The growing complexity of GNNs and size of real-w
 orld graphs have increased the memory requirements of GNN training and pop
 ular training platforms, like GPU, have memory capacity on the s...\n\n\nM
 ilan Shah, Reece Neff, and Michela Becchi (North Carolina State University
 )\n---------------------\nA Quantum Solver for Multidimensional Partial Di
 fferential Equations: Practical Case Studies\n\nQuantum computing is consi
 stently becoming transformational for computational problem-solving. This 
 capability appears particularly suited for numerical solution of multidime
 nsional partial-differential-equations (PDEs). Although many quantum techn
 iques are currently available for solving PDEs, thes...\n\n\nManu Chaudhar
 y (Illinois State University (ISU)) and Kareem El-Araby, Alvir Nobel, Ishr
 aq Islam, Manish Singh, Sunday Ogundele, Kieran Egan, Sneha Thomas, Vincen
 t Vordtriede, Devon Bontrager, Serom Kim, and Esam El-Araby (University of
  Kansas)\n---------------------\nHydraCache: LLM Inference Prefill Paralle
 lization Through Distributed Cache Blending\n\nThe prefill phase of large 
 language model (LLM) inference, where the input prompt is processed to gen
 erate a key-value (KV) cache, is a critical latency bottleneck for input s
 equences. Existing serving architectures face a trade-off: data parallelis
 m (DP) offers flexibility but cannot accelerate a s...\n\n\nAdib Rezaei Sh
 ahmirzadi (Virginia Tech), Shayan Shabihi (University of Maryland), Mona M
 oghadampanah (Virginia Tech), Furong Huang (University of Maryland), and D
 imitrios S. Nikolopoulos (Virginia Tech)\n---------------------\nWONDERS: 
 Integrating WOW, PONDER, and SCALE for Enhanced Scheduling Performance\n\n
 Recent scientific workflow management systems, such as Nextflow, put a hug
 e focus on portability of workflows. Portability encompasses replacing bot
 h the target infrastructure and the input dataset. The more portable syste
 ms become, the more the importance of automatic adaptation and optimizatio
 n in...\n\n\nFabian Lehmann (Humboldt-Universität zu Berlin); Jonathan Rau
 , Jonathan Bader, and Odej Kao (Technical University of Berlin); and Ulf L
 eser (Humboldt-Universität zu Berlin)\n---------------------\nJACC: Easy C
 PU/GPU Performance Portability for Scientific Applications in Julia\n\nOur
  JACC poster for SC25 presents a completely updated version of the Best Po
 ster finalist at SC24, showcasing the latest added features in the JACC li
 brary and ecosystem for productive scientific computing. First, we describ
 e the new and stable JACC API components: (i) a portable memory model, (ii
 )...\n\n\nWilliam Godoy, Pedro Valero-Lara, Philip Fackler, Keita Teranish
 i, and Jeffrey Vetter (Oak Ridge National Laboratory (ORNL)); Jhonny Gonza
 lez and Jose Gonzalez (The University of Texas at El Paso, Oak Ridge Natio
 nal Laboratory (ORNL)); and Alexis Huante (The University of Texas at Aust
 in, Oak Ridge National Laboratory (ORNL))\n---------------------\nMassivel
 y Parallel Bayesian Inference Framework for GPU Supercomputers: Applicatio
 n to Estimation of Coseismic Fault Slip\n\nWe present a massively parallel
  Bayesian inference framework for GPU supercomputers, demonstrated in cose
 ismic fault slip estimation. Bayesian inference, a robust method for inver
 se analysis, often relies on Monte Carlo sampling with over 100,000 forwar
 d simulations, making large-scale applications ...\n\n\nKai Nakao, Tsuyosh
 i Ichimura, and Kohei Fujita (The University of Tokyo)\n------------------
 ---\nAccelerating AI Co-Scientists with HPC Infrastructure\n\nWe present M
 OFAI, an agentic AI scientist coupled with high performance computing (HPC
 ) resources for the generation and property prediction of metal-organic fr
 ameworks (MOFs). MOFAI exhibits autonomous agents to enable tool-calling o
 f linker generation, MOF assembly, molecular dynamics simulation, ...\n\n\
 nSuryatejas Appana (University of California, Berkeley)\n-----------------
 ----\nAutoSlim: Intelligent Automata Graph Optimization for Efficient Acce
 leration\n\nModern high performance computing increasingly relies on sophi
 sticated graph-based models to represent and manipulate symbolic data. Fro
 m bioinformatics and cyber security to inference of the AI model and text 
 analytics, these applications often use directed graphs to capture complex
  dependencies an...\n\n\nTiffany Yu and Rasha Karakchi (University of Sout
 h Carolina)\n---------------------\nTowards Application Agnostic HPC Profi
 ling\n\nModern HPC systems generate large amounts of GPU and network telem
 etry, typically used for system health monitoring. At NERSC, we are develo
 ping a Performance API/UI that generates a job report card from this telem
 etry, providing an overview of performance characteristics. Using DCGM cou
 nters, we re...\n\n\nHari Teja Jajula (The University of Alabama, Lawrence
  Berkeley National Laboratory (LBNL)); Dhruva Kulkarni and Brian Austin (L
 awrence Berkeley National Laboratory (LBNL)); and Purushotham Bangalore (T
 he University of Alabama)\n---------------------\nLeveraging Large Languag
 e Models for Property Prediction in Polymorphic Organic Semiconductors\n\n
 Organic semiconductors (OSCs) are promising for next-generation electronic
 s, but polymorphism complicates accurate property prediction and makes tra
 ditional methods costly. We investigate transformer-based large language m
 odels (LLMs) for predicting energy gaps in polymorphic OSC crystals. A Peg
 asus...\n\n\nShreya Pagaria (Carnegie Mellon University, Pittsburgh Superc
 omputing Center) and Mei-Yu Wang, Dana O’Connor, Julian Uran, and Paola Bu
 itrago (Pittsburgh Supercomputing Center)\n---------------------\nEuropean
  Open Web Index: Large Complex Graph Visualization\n\nVisualization and pr
 ocessing of (extreme) large-scale networks is a challenging task due to un
 ique characteristics such as load imbalance, lack of locality, and access 
 irregularity. Considering the possibilities offered by recent supercomputi
 ng power, we have revised current algorithms suitable for ...\n\n\nPavlina
  Smolkova and Katerina Slaninova (IT4Innovations, VSB - Technical Universi
 ty of Ostrava)\n---------------------\nNumerical Investigation of Radiatio
 n Hydrodynamic Instabilities at Scale with FleCSI-HARD\n\nThis poster desc
 ribes the open-source massively parallel and portable radiation hydrodynam
 ics code FleCSI-HARD (Hydrodynamic And Radiative Diffusion) used to study 
 radiation hydrodynamics instabilities. FleCSI-HARD is based on FleCSI (Fle
 xible Computational Science Infrastructure) runtime which enab...\n\n\nMån
 s I. Andersson (KTH Royal Institute of Technology), Isaac C. Bannerman (Re
 nsselaer Polytechnic Institute), Moon B. Hazarika (University of Michigan)
 , Akshit Jariwala (The University of Texas at Austin), Jonathan Mathurin (
 Florida International University), Madela B. Quashie (Michigan State Unive
 rsity), and Julien Loiseau and Hyun Lim (Los Alamos National Laboratory (L
 ANL))\n---------------------\nScalable Execution Framework for R on Manyco
 re Systems\n\nRCOMPSs is a scalable execution framework that integrates th
 e R programming language with the COMPSs runtime to enable task-based para
 llel execution on manycore and distributed systems. RCOMPSs extends conven
 tional R workflows by allowing functions to be annotated as tasks, which t
 he runtime system ...\n\n\nXiran Zhang (King Abdullah University of Scienc
 e and Technology (KAUST)), Javier Conejero (Barcelona Supercomputing Cente
 r (BSC)), Sameh Abdulah (King Abdullah University of Science and Technolog
 y (KAUST)), Jorge Ejarque (Barcelona Supercomputing Center (BSC)), Ying Su
 n (King Abdullah University of Science and Technology (KAUST)), Rosa M. Ba
 dia (Barcelona Supercomputing Center (BSC)), and David E. Keyes and Marc G
 . Genton (King Abdullah University of Science and Technology (KAUST))\n---
 ------------------\nAdvancing EEG Signal Analysis with Quantum Machine Lea
 rning\n\nElectroencephalography (EEG) is widely used in brain–computer int
 erfaces, but movement-related signals are weak, variable, and often buried
  in noise. Classical pipelines, such as Random Forests trained on PCA+CSP 
 features, work fairly well but can miss cross-channel patterns. Quantum ma
 chine l...\n\n\nStephanie Murray (University of Washington Bothell, Univer
 sity of Hawaii at Manoa) and Erika Parsons (University of Washington Bothe
 ll)\n---------------------\nEnergy-Efficient Multimodal LLM Inference: Sta
 ge-Level Characterization and Input-Aware Controls​\n\nMultimodal large la
 nguage models (MLLMs) extend text-only LLMs with image and video encoders,
  enabling new capabilities but introducing high and poorly understood ener
 gy costs. This work characterizes the energy footprint of MLLM inference a
 t the stage level, decomposing serving into vision encoding...\n\n\nMona M
 oghadampanah, Adib Rezaei Shahmirzadi, and Dimitrios S. Nikolopoulos (Virg
 inia Tech)\n---------------------\nCUR-MoE: Portable Mixture-of-Experts wi
 th Interpretable High-Ratio Compression\n\nMixture-of-experts (MoE) archit
 ectures enable trillion-parameter models but face prohibitive memory scali
 ng, limited compression interpretability, and vendor-specific implementati
 ons hindering heterogeneous HPC deployment.\n\nWe present the first Julia-
 based MoE framework introducing CUR decomposition...\n\n\nRitesh Bhirud (U
 niversity of Massachusetts Amherst, Massachusetts Institute of Technology 
 (MIT))\n---------------------\nUnraveling Distant Galaxies: Analyzing IFU 
 Data with Parsl and Academy\n\nIntegral field spectroscopy is a powerful t
 echnique in observational astrophysics enabling the study of spatially-com
 plex objects like distant strongly-lensed galaxies. Integral field units (
 IFUs) are an increasingly common addition to many powerful ground- and spa
 ce-based observatories, which motiv...\n\n\nDaniel Babnigg (University of 
 Chicago)\n---------------------\nBuilding the Foundation for Machine Learn
 ing-Based Mars Weather Forecasting\n\nMars is a leading target for human e
 xploration, yet its weather remains difficult to predict due to phenomena 
 such as global dust storms. While Earth forecasting has advanced through m
 achine learning (ML), Mars lacks comparable systems. This work investigate
 s whether Microsoft’s Aurora, a stat...\n\n\nMohammad Altiwainy (Wayne Sta
 te University)\n---------------------\nScaling Singular Values Beyond GPU 
 Memory Limits: Out-of-Core, GPU-Accelerated, and Unified Across Data Preci
 sion and Hardware\n\nWe present a unified, out-of-core, GPU-accelerated si
 ngular value solver that achieves performance portability across diverse h
 ardware platforms and data precisions for datasets exceeding GPU memory. T
 he singular value decomposition (SVD) is fundamental for processing large-
 scale datasets, yet the d...\n\n\nEvelyne Ringoot (Massachusetts Institute
  of Technology (MIT))\n---------------------\nCan Long-Haul RDMA Benefit F
 ederated Learning?\n\nFederated learning (FL) has emerged as a promising p
 aradigm for privacy-preserving distributed training. However, its performa
 nce is often hindered by communication bottlenecks, especially over long-d
 istance networks. In this work, we investigate the effectiveness of long-h
 aul remote direct memory a...\n\n\nZhonghao Chen, Yuke Li, Duo Zhang, and 
 Xiaoyi Lu (University of California, Merced)\n---------------------\nThe I
 mpact of Maximum Vector Length on Cache Management Techniques in RISC-V Ve
 ctor Extension\n\nIn recent years, the RISC-V vector extension (RVV) has a
 ttracted increasing attention. The RVV allows programs to be executed on p
 rocessors with various maximum vector lengths (MVLs). Consequently, even w
 hen running the same program, the memory access pattern may vary depending
  on the MVL of the pro...\n\n\nShunya Nomura (Tohoku University); Jiaheng 
 Liu (RIKEN Center for Computational Science (R-CCS)); Keichi Takahashi (Th
 e University of Osaka, Tohoku University); and Hiroyuki Takizawa (Tohoku U
 niversity)\n---------------------\nDistributed 3D Gaussian Splatting for H
 igh-Resolution Isosurface Visualization\n\n3D Gaussian Splatting (3D-GS) h
 as recently emerged as a powerful technique for real-time, photorealistic 
 rendering by optimizing anisotropic Gaussian primitives from view-dependen
 t images. While 3D-GS has been extended to scientific visualization, prior
  work remains limited to single-GPU settings, r...\n\n\nMengjiao Han (Argo
 nne National Laboratory (ANL)); Andres Sewell (Utah State University); Jos
 eph Insley and Janet Knowles (Argonne National Laboratory (ANL)); Victor A
 . Mateevitsi and Michael E. Papka (Argonne National Laboratory (ANL), Univ
 ersity of Illinois Chicago); Steve Petruzza (Utah State University); and S
 ilvio Rizzi (Argonne National Laboratory (ANL))\n---------------------\nFo
 rward Error Bounds and Efficient Algorithms for Computing a Tensor Times M
 atrix Chain in Low Precision on GPUs\n\nMany tensor processing algorithms 
 require computing a tensor times matrix chain (TTMc) operation, and this o
 peration is frequently the bottleneck in such algorithms. This work develo
 ps strategies for accelerating a TTMc using low-precision hardware. \n\nWe
  present a novel scheme for scaling the TTMc o...\n\n\nJulian Bellavita (C
 ornell University, Oak Ridge National Laboratory (ORNL)) and Piyush Sao an
 d Ramakrishnan Kannan (Oak Ridge National Laboratory (ORNL))\n------------
 ---------\nBridging the Quantum Coding Gap: Instruction-Tuned LLMs for Qis
 kit\n\nLarge language models (LLMs) have advanced code generation ability 
 across many domains, but often struggle with quantum code due to limited d
 omain-specific data and inherent domain complexity. To address this issue,
  we focus on the Qiskit framework and fine-tune pretrained LLMs using quan
 tum code fr...\n\n\nSixu Chen, Yuqi Zhang, and Qiang Guan (Kent State Univ
 ersity)\n---------------------\nLearning To Select Scheduling Algorithms i
 n OpenMP\n\nScientific and data science applications demand increasing com
 putational performance, requiring effective scheduling and load balancing 
 on high performance computing (HPC) systems. While OpenMP libraries such a
 s LB4OMP provide several scheduling algorithms, selecting the best one for
  a given applica...\n\n\nJonas H. Müller Korndörfer (University of Bern, U
 niversity of Basel); Ali Mohammed and Ahmed Eleliemy (HPE HPC/AI EMEA Lab)
 ; Quentin Guilloteau (Inria); and Reto Krummenacher and Florina Ciorba (Un
 iversity of Basel)\n---------------------\nEvaluating the Power-Monitoring
  Capabilities of Aurora\n\nExascale systems like Aurora push performance b
 ounds but they draw tens of megawatts, making precise, low-overhead power 
 monitoring essential for efficiency and cost control. We present an ongoin
 g evaluation of the two primary power-monitoring interfaces on Aurora, qua
 ntifying accuracy and temporal ...\n\n\nPrecious Eyabi (Argonne National L
 aboratory (ANL))\n---------------------\nShortcut Mixup Policy: Toward Imp
 roving Robustness and Speed in Goal-Conditioned RL\n\nNeural networks trai
 ned on large datasets can be effective policies for the control of robotic
  manipulators. Using self-supervised learning, these networks can achieve 
 near-perfect success rates on complex pick-and-place-style tasks. However,
  the speed of task completion is often a barrier to making...\n\n\nMatthew
  Hyatt (Loyola University Chicago, Argonne National Laboratory (ANL)); Yas
 sir Atlas, Hal Brynteson, Diego Roa Perdomo, Athena Angara, Mengjiao Han, 
 Joseph Insley, Janet Knowles, Yongho Kim, Victor Mateevitsi, Michael Papka
 , and Silvio Rizzi (Argonne National Laboratory (ANL)); George Thiruvathuk
 al (Loyola University Chicago, Argonne National Laboratory (ANL)); and Nic
 ola Ferrier (Argonne National Laboratory (ANL))\n---------------------\nPr
 oductive Scalable Distributed Task Scheduling Using an MPI-based Backend f
 or Dagger\n\nDistributed computing frameworks are vital for managing compl
 ex workloads in high-performance computing and scientific research. Julia'
 s Dagger.jl supports task-based parallelism using TCP communication, suita
 ble for cloud and local environments. However, TCP limits performance on m
 odern HPC systems...\n\n\nYan Guimarães (University of Brasilia)\n--------
 -------------\nUnified Performance Modeling Stack for Distributed GPU Appl
 ications: Complementing Analytical Insights with Machine Learning\n\nModer
 n HPC applications increasingly use GPUs to solve larger problems with hig
 her accuracy and speed. However, committing resources to these large-scale
  systems is often costly and time-consuming. Hence, performance modeling e
 nables developers to estimate runtime, analyze scalability, and identify .
 ..\n\n\nUrvij Saroliya (Technical University of Munich)\n-----------------
 ----\nPerformance Engineering of Scientific Applications with MVAPICH and 
 TAU Using Emerging Communication Primitives\n\nWe propose a co-design appr
 oach that integrates two powerful tools—MVAPICH and TAU—to demonstrate the
  new possibilities for performance-guided control and optimization for two
  large-scale applications—AWP-ODC and heFFTe. AWP-ODC is a highly scalable
  parallel finite-difference appli...\n\n\nDhabaleswar K. (DK) Panda (The O
 hio State University); Sameer Shende (University of Oregon; ParaTools, Inc
 .); Ahmad Abdelfattah (University of Tennessee, Knoxville); and Yifeng Cui
  (San Diego Supercomputer Center (SDSC))\n---------------------\nChatHPC: 
 Building the Foundations for a Productive and Trustworthy AI-Assisted HPC 
 Ecosystem\n\nChatHPC democratizes large language models for the high perfo
 rmance computing (HPC) community by providing the infrastructure, ecosyste
 m, and knowledge needed to apply modern generative AI technologies to rapi
 dly create specific capabilities for critical HPC components while using r
 elatively modest ...\n\n\nPedro Valero-Lara, Aaron Young, Mohammad Alaul H
 aque Monil, Swaroop Pophale, Zheming Jin, Jeffrey S. Vetter, Keita Teranis
 hi, and William F. Godoy (Oak Ridge National Laboratory (ORNL))\n---------
 ------------\nWafer-Scale Simulation of Mutator Allele Dynamics in Large A
 sexual Populations\n\nWithin evolving microbial populations, genes that el
 evate mutation rate impose a fundamental trade-off: on one hand, increasin
 g harmful mutations among offspring, but, on the other, allowing more oppo
 rtunities for rare beneficial mutations. Existing single-CPU agent-based s
 imulation work suggests th...\n\n\nMatthew Andres Moreno (University of Mi
 chigan), Emily Dolson (Michigan State University), and Luis Zaman (Univers
 ity of Michigan)\n---------------------\nMPI-SGX: Enabling Confidential Co
 mputing for MPI Parallel Applications with Intel SGX Technology\n\nBig dat
 a and deep learning workloads often require handling sensitive data, but s
 ecurity mechanisms in current supercomputers mainly protect against extern
 al threats, leaving risks of insider leakage. As a result, supercomputers 
 remain unsuitable for confidential applications. To address this challe...
 \n\n\nKota Shimojima (The University of Electro-Communications, RIKEN Cent
 er for Computational Science (R-CCS)); Hayato Yamaki and Hiroki Honda (The
  University of Electro-Communications); Shinichiro Matsuo (Georgetown Univ
 ersity); Atsuko Takefusa (National Institute of Informatics, Japan; RIKEN 
 Center for Computational Science (R-CCS)); and Shinobu Miwa (The Universit
 y of Electro-Communications, RIKEN Center for Computational Science (R-CCS
 ))\n---------------------\nUnderstanding LLM Behavior on HPC Data via Mech
 anistic Interpretability\n\nLarge language models (LLMs) are increasingly 
 used in HPC for tasks like code generation and analysis, but their interna
 l reasoning remains opaque. To address this, we study three tasks—OpenMP c
 ode completion, data race detection, and OMP code generation—using mechani
 stic interpretabilit...\n\n\nMd Mahbubur Rahman (Iowa State University), A
 rjun Guha (Northeastern University), and Harshitha Menon (Lawrence Livermo
 re National Laboratory (LLNL))\n---------------------\nAnalyzing Dataset P
 opularity for Optimizing In-Network Storage\n\nIn high energy physics (HEP
 ), large-scale experiments produce enormous data volumes that are distribu
 ted across global storage systems. To reduce redundant transfers and impro
 ve efficiency, disk caching systems such as XCache are deployed, but their
  effectiveness depends on good caching policy. Our ...\n\n\nGunwoo Kim (Un
 iversity of California, Davis) and Alex Sim and Kesheng Wu (ESnet; Lawrenc
 e Berkeley National Laboratory (LBNL))\n---------------------\nMojo: Pytho
 n-Like MLIR-Based GPU Portable Science Kernels\n\nThis work investigates M
 ojo, a new MLIR-based language that combines Python-like syntax with porta
 ble, low-level GPU programming capabilities. We compare the performance of
  the Mojo portable GPU kernels against vendor-specific C++ NVIDIA CUDA and
  AMD HIP implementations on four representative scient...\n\n\nTatiana Mel
 nichenko (University of Tennessee, Knoxville; Oak Ridge National Laborator
 y (ORNL))\n---------------------\nHigh Performance Batch SVD Using GPUs\n\
 nWe consider the problem of computing the singular value decomposition (SV
 D) of many relatively small matrices using GPUs. This is an essential comp
 onent in various scientific applications, including computational chemistr
 y, low-rank approximations, and others. Our approach is based on the paral
 lel o...\n\n\nAhmad Abdelfattah (University of Tennessee, Knoxville)\n----
 -----------------\nPractical Viability of Translating Legacy Fortran Code 
 to C++ Using Large Language Models\n\nIn this study, we discuss the practi
 cality and limitations of using current large language models (LLMs) for a
 utomatically translating Fortran legacy codes to C++, so that legacy codes
  written in Fortran can be modernized to exploit the performance and featu
 res available only in C++. Moreover, we in...\n\n\nRen Imai and Masatoshi 
 Kawai (Tohoku University), Keichi Takahashi (The University of Osaka), and
  Hiroyuki Takizawa (Tohoku University)\n---------------------\nEnabling Ef
 ficient Runtime Data Analysis to a Crystal Deformation Simulation\n\nExasc
 ale simulations generate massive data volumes that strain I/O and post-hoc
  analysis. We integrate the Damaris in situ middleware into Coddex, a crys
 tal deformation code, to offload data movement and analysis to dedicated p
 rocesses, enabling runtime extraction of key diagnostics without writing .
 ..\n\n\nArthur Jaquard (Inria)\n---------------------\nUnderstanding Commu
 nication Bottlenecks in Multi-Node LLM Inference\n\nAs large language mode
 ls (LLMs) grow in parameter count, efficient generation requires inference
  to scale beyond a single node. Current approaches use tensor parallelism 
 (TP) or pipeline parallelism (PP), but TP incurs high communication volume
 , while PP suffers from pipeline bubbles and is unsuitab...\n\n\nPrajwal S
 inghania (University of Maryland); Siddharth Singh (University of Maryland
 , NVIDIA Corporation); Lannie Dalton Hough and Ishan Revankar (University 
 of Maryland); Harshitha Menon and Charles Jekel (Lawrence Livermore Nation
 al Laboratory (LLNL)); and Abhinav Bhatele (University of Maryland)\n-----
 ----------------\nMassively Parallel GPU Rasterizer for Next-Generation Co
 mputational Lithography\n\nIn electronic design automation (EDA), traditio
 nal rasterization algorithms suffer from poor speedup and accuracy when ma
 naging large and complex semiconductor designs, limiting efficiency in opt
 ical proximity correction (OPC) processes. To overcome these challenges, w
 e developed a GPU-based rasteri...\n\n\nLoay Hegazy, Mohamed Taher, and Sh
 erif Hammouda (Siemens EDA)\n---------------------\nWhen Label Propagation
  Outperforms BFS in Breadth-First Graph Traversal\n\nWe tackle the challen
 ge of breadth-first traversal (BFT) on sparse graphs with a high number of
  connected components. We propose a novel distributed-memory parallel algo
 rithm that uses the label propagation (LP) algorithm to perform BFT on all
  connected components of the graph simultaneously. In syn...\n\n\nKalsuda 
 Lapborisuth and Srinivas Aluru (Georgia Institute of Technology)\n--------
 -------------\nShipping HPC Ecosystems Across Platforms: Portable and Comp
 osable HPC Clusters as Code\n\nHigh performance computing systems require 
 intricate platform-specific stacks and configurations, which poses a chall
 enge for reproducing the same HPC ecosystem on a different platform, a key
  requirement for geo-redundancy, business continuity, and urgent computing
 . We present a method for declarati...\n\n\nGerman Felipe Giraldo Villa, T
 héo Grivel, George Ioannidis, Edita Kizinevic, Carolina Lindqvist, Nicolas
  Litchinko, Pablo Llopis, Antonio Javier Russo, and Gilles Fourestey (Écol
 e Polytechnique Fédérale de Lausanne)\n---------------------\nParaViz3D: M
 PI Trace Visualization with 3D Video\n\nConventional MPI trace visualizati
 ons become unwieldy as the number of processes increases. Identifying glob
 al patterns is challenging, and both the underlying machine topology and p
 arallel program structure are often obscured by the rigid two-dimensional 
 rank-time graphs. To address these limitatio...\n\n\nJean-Yves Verhaeghe, 
 Georg Hager, and Ayesha Afzal (Friedrich-Alexander-Universität Erlangen-Nü
 rnberg, Erlangen National High Performance Computing Center)\n------------
 ---------\nEvaluating the Usage of Python Libraries on a Production Superc
 omputer\n\nAdvances in artificial intelligence (AI) and machine learning (
 ML) are reshaping scientific computing and influencing programming practic
 es on high performance computing (HPC) systems. We analyze Python library 
 usage on the Polaris supercomputer to understand adoption patterns in mode
 ling, simulatio...\n\n\nThomas Papka (Loyola University Chicago, Argonne N
 ational Laboratory (ANL))\n---------------------\nDetecting Silent Data Co
 rruption in Sparse Matrices Using Hardware Performance Counters\n\nHigh pe
 rformance computing (HPC) systems frequently execute large-scale sparse ma
 trix computations in scientific and engineering domains. These workloads a
 re susceptible to silent data corruptions (SDCs)—undetected faults that ca
 n alter results without triggering errors—posing a signific...\n\n\nMinseo
 p Choi, Orlando Arias, and Seung Woo Son (University of Massachusetts Lowe
 ll)\n---------------------\nCROSS-HPC System Bayesian Optimization with Ad
 aptive Transfer\n\nThis paper introduces CROSS BOAT (Cross HPC System Baye
 sian Optimization with Adaptive Transfer), a novel method for efficient pa
 rameter tuning in high performance computing (HPC) systems. Optimizing the
  many configurable parameters in HPC environments usually requires costly 
 evaluations on each tar...\n\n\nAbrar Hossain and Kishwar Ahmed (Universit
 y of Toledo)\n---------------------\nDiOMP-Offloading: Portable OpenMP Off
 loading for Distributed Heterogeneous Systems\n\nAs heterogeneous supercom
 puting becomes mainstream, traditional hybrid models such as MPI+OpenMP in
 creasingly struggle to coordinate and manage GPU memory while maintaining 
 portable performance. \n\nThis work introduces DiOMP-Offloading, a framewo
 rk that unifies OpenMP target offloading with a PGAS mo...\n\n\nBaodi Shan
  (Stony Brook University), Mauricio Araya-Polo (TotalEnergies), and Barbar
 a Chapman (Stony Brook University)\n---------------------\nSync-Free GPU P
 arallelization of Sparse Kernels from Sequential Python Code\n\nSparse mat
 rix kernels such as SpMV, SpTRSV, and Gauss-Seidel are critical in scienti
 fic computing, AI, and engineering, but they remain difficult to paralleli
 ze due to irregular memory access patterns. Traditional compiler technique
 s assume affine array accesses, which do not hold in sparse formats ...\n\
 n\nMalko-Bani Somo (McMaster University)\n---------------------\nUnmasking
  Performance Variability in GPU Codes on Supercomputers\n\nPerformance var
 iability is often a critical issue on GPU-accelerated systems, undermining
  efficiency and reproducibility. Since large-scale investigations of perfo
 rmance variability on GPU clusters are lacking, we set up a longitudinal e
 xperiment on Perlmutter and Frontier. We benchmark representati...\n\n\nCu
 nyang Wei and Keshav Pradeep (University of Maryland, College Park) and Ab
 hinav Bhatele (University of Maryland)\n---------------------\nReal-Time M
 L-Based Defense Against Malicious Payload in Reconfigurable Embedded Syste
 ms\n\nField-programmable gate arrays (FPGAs) in reconfigurable systems fac
 e escalating security threats from malicious bitstreams capable of causing
  denial-of-service, data leakage, or covert operations. Traditional detect
 ion methods often require source code or netlists, limiting their applicab
 ility for ...\n\n\nRye Stahle-Smith and Rasha Karakchi (University of Sout
 h Carolina)\n---------------------\nRange Search on Heterogeneous Systems 
 with Processing-in-Memory Architecture\n\nThe growing volume of data in hi
 gh performance computing (HPC) has made spatial query processing increasin
 gly challenging due to high data transfer costs and limited memory bandwid
 th. To address these bottlenecks and reduce energy wasted on data movement
 , this work explores processing-in-memory (PIM...\n\n\nTasmia Jannat and S
 atish Puri (Missouri University of Science and Technology) and Michael Gow
 anlock (Northern Arizona University)\n---------------------\nA Scalability
  Study of Quantum Algorithms for Dimensionality Reduction of Multidimensio
 nal Data\n\nQuantum computing promises exponential speedups over classical
  computing by leveraging quantum-mechanical properties like superposition 
 and entanglement. As quantum algorithms grow in complexity, classical simu
 lation remains essential for evaluating correctness, scalability, and reso
 urce demands. Th...\n\n\nKareem El-Araby (University of Kansas); Thom Popo
 vic (Lawrence Berkeley National Laboratory (LBNL)); Alvir Nobel and Sunday
  Ogundele (University of Kansas); Katherine Klymko, Daan Camps, and Anasta
 siia Butko (Lawrence Berkeley National Laboratory (LBNL)); and Esam El-Ara
 by (University of Kansas)\n---------------------\nAccelerating Linear Solv
 e with Mixed Precision Nested Recursive Subdivision on AI Hardware\n\nThe 
 Cholesky decomposition is a critical performance bottleneck in engineering
  simulations. To accelerate these simulations, we present a novel, nested 
 recursive Cholesky algorithm implemented in Julia. The algorithm restructu
 res the problem into recursive TRSM (triangular solve) and SYRK (symmetric
 ...\n\n\nVicki Carrica (Massachusetts Institute of Technology (MIT))\n----
 -----------------\nMulti-GPU Implementation and Roofline Analysis of a Num
 erical Global Ocean Model\n\nNumerical ocean models are essential tools fo
 r climate prediction and marine resource studies, requiring high resolutio
 n and realistic physical processes. We developed the global ocean model CO
 CO and implemented it on GPUs using an OpenACC directive-based approach, w
 hile maintaining compatibility wi...\n\n\nTakateru Yamagishi (Research Org
 anization for Information Science and Technology); Masao Kurogi and Takao 
 Kawasaki (Japan Agency for Marine-Earth Science and Technology); Yoshimasa
  Matsumura (National Institute for Environmental Studies); and Hiroyasu Ha
 sumi (Atmosphere and Ocean Research Institute, The University of Tokyo)\n-
 --------------------\nDiffPro: Joint Timestep and Layer-Wise Precision Opt
 imization for Efficient Diffusion Inference\n\nDiffPro is a simple framewo
 rk to speed up and shrink diffusion models while preserving image quality.
  It combines layer-wise quantization, guided by a manifold-based sensitivi
 ty check with adaptive timestep selection. Compared with quantization-only
  or sampling-only baselines, this joint strategy yi...\n\n\nFarhana Amin (
 Virginia Tech), Kanchon Gharami (Embry-Riddle Aeronautical University), an
 d Dimitrios S. Nikolopoulos (Virginia Tech)\n---------------------\nCan Lo
 ssy Compression Benefit NVMe-Based I/O?\n\nLossy compression is widely use
 d to reduce storage costs and I/O demands, especially on SCSI-HDDs. Howeve
 r, its benefits diminish on NVMe-SSDs, where compression and decompression
  runtimes often exceed raw I/O speed. To address this, we conduct a detail
 ed study of compression runtimes, control metho...\n\n\nDarren Ng and Duo 
 Zhang (University of California, Merced); Sheng Di (Argonne National Labor
 atory (ANL)); Zhaorui Zhang (The Hong Kong Polytechnic University); and Xi
 aoyi Lu (University of California, Merced)\n---------------------\nIncineR
 ate: Multi-Modal FPGA Accelerator for SCNNs\n\nSpiking neural networks (SN
 Ns) are a promising alternative to conventional artificial neural networks
  (ANNs) due to their biological interpretability and capability to exploit
  sparse computation. Specialized hardware for SNNs has advantages over gen
 eral-purpose devices in terms of power and performa...\n\n\nBjörn A. Lindq
 vist and Artur Podobas (KTH Royal Institute of Technology)\n--------------
 -------\nHarmony: Converged Supercomputer Scratch and Archival Filesystems
 \n\nIn high-performance computing, scratch storage holds intermediate data
  while archival storage holds long-term data. These distinct objectives le
 ad to separate implementations, implying additional costs and requiring ex
 plicit data transfers. This separation is inefficient, reduces reliability
 , and in...\n\n\nJake Carroll (The University of Queensland)\n------------
 ---------\nOptimizing Collectives with Large Payloads on GPU-Based Superco
 mputers\n\nWe evaluate the current state of collective communication on GP
 U-based supercomputers for large language model (LLM) training at scale. E
 xisting libraries such as RCCL and Cray-MPICH exhibit critical limitations
  on systems such as Frontier—Cray-MPICH underutilizes network and compute 
 resources...\n\n\nSiddharth Singh (NVIDIA Corporation, University of Maryl
 and); Mahua Singh (IIT Guwahati); and Keshav Pradeep and Abhinav Bhatele (
 University of Maryland)\n---------------------\nCATIOS: Time-Resolved I/O-
 Aware Job Scheduling for HPC Systems\n\nHPC workloads are increasingly dat
 a-intensive, with contention on shared storage emerging as a primary bottl
 eneck. Existing I/O-aware job schedulers rely on static bandwidth assumpti
 ons that overlook time-varying I/O behavior, leading to inefficient utiliz
 ation and unpredictability.\n\nThis work intro...\n\n\nYuTsen Tseng (Tohok
 u University, Graduate School of Information Sciences); Masatoshi Kawai (T
 ohoku University); Keichi Takahashi (University of Osaka); and Hiroyuki Ta
 kizawa (Tohoku University)\n---------------------\nDivergence Prediction S
 ystem for CFD Simulations\n\nComputational fluid dynamics (CFD) simulation
 s are essential tools for analyzing complex flow phenomena in engineering 
 and scientific research. These simulations are typically formulated based 
 on the Navier-Stokes equations, which govern the motion of incompressible 
 fluids, and the pressure field is...\n\n\nTakashi Soga (The University of 
 Osaka), Takanori Uchida (Kyushu University), and Susumu Date (The Universi
 ty of Osaka)\n---------------------\nAn Approach for Correlating Compiler 
 Optimizations with Runtime Performance\n\nPerformance-portability librarie
 s such as RAJA enable single-source applications to run on diverse archite
 ctures, but performance often depends on compiler decisions that are hard 
 to observe. Existing tools either show compiler activity without runtime c
 ontext or runtime performance without compiler...\n\n\nBefikir Bogale (Uni
 versity of Tennessee, Knoxville); Olga Pearce (Lawrence Livermore National
  Laboratory); Tom Scogland (Lawrence Livermore National Laboratory (LLNL))
 ; and Michela Taufer (University of Tennessee, Knoxville)\n---------------
 ------\nTime-Stepping Hamiltonian Simulation for Solving Nonlinear PDEs vi
 a a Quantum-Classical Hybrid Approach\n\nThis work presents a time-steppin
 g Hamiltonian simulation framework for nonlinear PDEs on a hybrid quantum–
 classical approach. Using warped phase transform (WPT)–based Schrödingeriz
 ation, spatial discretizations are reformulated as Hermitian/anti-Hermitia
 n operators for Schrödinger-type ...\n\n\nSangwon Kim and Junya Onishi (RI
 KEN Center for Computational Science (R-CCS)); Ayato Takii (Kobe Universit
 y, Japan); Younghwa Cho (Hokkaido University, Japan); and Tsubokura Makoto
  (RIKEN Center for Computational Science (R-CCS); Kobe University, Japan)\
 n---------------------\nInference-as-a-Service Prototype at NERSC\n\nThe i
 ncreasing scale and complexity of scientific experiments has led to a grow
 ing need for efficient and scalable machine learning model inference servi
 ng systems. High-energy physics experiments and simulations of complex cli
 mate models involve petabytes of data and massive amounts of computationa.
 ..\n\n\nColin Thomas (University of Notre Dame); Po-Han Huang (Georgia Ins
 titute of Technology); Hilary Utaegbulam (University of Rochester); Johann
 es Blaschke (ESnet; Lawrence Berkeley National Laboratory (LBNL)); Bruno C
 oimbra (Fermi National Laboratory); Pengfei Ding, Xiangyang Ju, and Andrew
  Naylor (ESnet; Lawrence Berkeley National Laboratory (LBNL)); and Michael
  Wang (Fermi National Laboratory)\n---------------------\nClassifying Perf
 ormance Bounds Using Machine Learning\n\nTraditional performance analysis 
 tools, such as the roofline model, require visual interpretation to determ
 ine performance bounds. For CPUs which have complex cache hierarchies and 
 front-end out-of-order capabilities—that is, the CPUs we use for high perf
 ormance computing—accurately iden...\n\n\nLewis Littman and Tom Deakin (Un
 iversity of Bristol)\n---------------------\nOrchid: Towards Heterogeneous
  Batched Eigenvalue Solvers\n\nThere is a growing need for the efficient s
 olution of many small eigenvalue problems (up to N = 1500) that arise in e
 merging scientific applications. These small-to-medium sized problems pres
 ent unique computational challenges, particularly when thousands or millio
 ns of such problems must be solved ...\n\n\nMatthew Chung (University of C
 alifornia, Riverside; Oak Ridge National Laboratory (ORNL))\n-------------
 --------\nExploring Fine-Grained Parallelism in Data-Flow Runtime Systems 
 on Many-Core Systems\n\nHigh synchronization overhead in frameworks like G
 NU OpenMP impedes fine-grained task parallelism on many-core architectures
 . We introduce three advances to GNU OpenMP: a lock-less concurrent queue 
 (XQueue), a scalable distributed tree barrier, and two NUMA-aware, lock-le
 ss load-balancing strategies...\n\n\nWenyi Wang and Maxime Gonthier (Unive
 rsity of Chicago), Haibin Lai (Southern University of Science and Technolo
 gy), Poornima Nookala (Intel Corporation), Haochen Pan and Ian Foster (Uni
 versity of Chicago), Ioan Raicu (Illinois Institute of Technology), and Ky
 le Chard (University of Chicago)\n---------------------\nChameleon Concier
 ge: Retrieval-Augmented Generation (RAG) To Enhance Open Testbed Documenta
 tion\n\nResearchers in high performance computing (HPC) and cloud environm
 ents encounter disparate sources of documentation and difficulties finding
  accurate information. This can cause inefficiency, increase the reliance 
 on support teams, and change the focus of the researcher from the main exp
 eriment. To ...\n\n\nSaieda Ali Zada (University of Delaware)\n-----------
 ----------\nIntelligent Surrogates Pay Attention to Data, Improving Multi-
 Objective HPC Optimization\n\nHigh performance computing (HPC) schedulers 
 must balance runtime and power. We present a surrogate-assisted multi-obje
 ctive Bayesian optimization (MOBO) framework using TabNet regressors and m
 odels trained on attention-based embeddings, coupled with active-learning 
 sample selection. The surrogates p...\n\n\nAshna Nawar Ahmed (Texas State 
 University, Oak Ridge National Laboratory (ORNL)); Banooqa Banday (Texas S
 tate University); Terry Jones (Oak Ridge National Laboratory (ORNL)); and 
 Tanzima Z. Islam (Texas State University)\n\nTag: Research & ACM SRC Poste
 rs\n\nRegistration Category: Technical Program Reg Pass\n\nSession Chairs:
  Kento Sato (RIKEN Center for Computational Science (R-CCS)); Chris Schlip
 alius (Pawsey Supercomputing Research Centre; Commonwealth Scientific and 
 Industrial Research Organisation (CSIRO), Australia); and Anja Gerbes (Geo
 rg-August-Universität Göttingen)
END:VEVENT
END:VCALENDAR
