Close

Session

Event Type
Research and ACM SRC Posters
TimeTuesday, 18 November 20255:15pm - 7:00pm CST
LocationSecond Floor Atrium
Tags
Research & ACM SRC Posters
Registration Categories
TP
Presentations
Scaling Singular Values Beyond GPU Memory Limits: Out-of-Core, GPU-Accelerated, and Unified Across Data Precision and Hardware
Echoes of Earth: Building an Autonomous Environmental Lab for Acoustic Sensing
Scalable Execution Framework for R on Manycore Systems
Sync-Free GPU Parallelization of Sparse Kernels from Sequential Python Code
Facilitating Mixed Python-Fortran HPC Codes: 4D Drift-Kinetic Simulations with Pyccel
Algorithms and Applications of Dynamic Network Analysis Using CANDY
When Label Propagation Outperforms BFS in Breadth-First Graph Traversal
A Formal Characterization of Non-Monotonicity in Tensor Cores
Towards a GPU-Accelerated Web-Based Graph Rendering Framework for Large-Scale Protein Networks
Unmasking Performance Variability in GPU Codes on Supercomputers
AutoSlim: Intelligent Automata Graph Optimization for Efficient Acceleration
Shipping HPC Ecosystems Across Platforms: Portable and Composable HPC Clusters as Code
Chameleon Concierge: Retrieval-Augmented Generation (RAG) To Enhance Open Testbed Documentation
Explicit Low-Order Finite-Element Wave Simulation Accelerated with Variable-Precision Computing Using INT8 Tensor Cores
VaultX Merge: Breaking Memory Barriers in Proof-of-Space Plot Generation
Optimizing Collectives with Large Payloads on GPU-Based Supercomputers
WiCAT: Reducing Congestion at Wireless Interfaces in Heterogeneous Architectures
Intelligent Surrogates Pay Attention to Data, Improving Multi-Objective HPC Optimization
CIRE: LLVM Analysis for Floating-Point Rounding Error Affected by Precision and Optimizations
GNNs on Evolving Graphs: A Benchmark of Incremental Updates and Meta-Learning Approaches
An Agent-Based Viral Venture: Adaptive Tool Selection for Scalable Genomics
PhySiViT: A Physics Simulation Vision Transformer
Enabling Efficient Runtime Data Analysis to a Crystal Deformation Simulation
Compute System Simulator: Modeling the Impact of Allocation Policy and Hardware Reliability on HPC Cloud Resource Utilization
Divergence Prediction System for CFD Simulations
JACC: Easy CPU/GPU Performance Portability for Scientific Applications in Julia
Massively Parallel Bayesian Inference Framework for GPU Supercomputers: Application to Estimation of Coseismic Fault Slip
Optimizing Task-Driven Offloading in LLVM
HydraCache: LLM Inference Prefill Parallelization Through Distributed Cache Blending
MPI-SGX: Enabling Confidential Computing for MPI Parallel Applications with Intel SGX Technology
Heterogeneity-Aware Task Allocation for Modern HPC Systems
Productive Scalable Distributed Task Scheduling Using an MPI-based Backend for Dagger
Optimizing the GPU All-Reduce Using Multiple Processes Per GPU
DiOMP-Offloading: Portable OpenMP Offloading for Distributed Heterogeneous Systems
CATIOS: Time-Resolved I/O-Aware Job Scheduling for HPC Systems
Can Long-Haul RDMA Benefit Federated Learning?
Accelerating AI Co-Scientists with HPC Infrastructure
Forward Error Bounds and Efficient Algorithms for Computing a Tensor Times Matrix Chain in Low Precision on GPUs
Between the NIC and a Hard Place: Evaluating 400 Gb/s Ethernet for HPC Data Transfers
Julia with Intelligent Runtime for Heterogeneous Computing
WONDERS: Integrating WOW, PONDER, and SCALE for Enhanced Scheduling Performance
Evaluating LiDAR Compression for 3D Semantic Segmentation in Diverse Off-Road Environments on GOOSE Dataset
Novel Graph Alignment Algorithms for Identifying Non-Determinism in Large-Scale Simulations
Accelerating Linear Solve with Mixed Precision Nested Recursive Subdivision on AI Hardware
Practical Viability of Translating Legacy Fortran Code to C++ Using Large Language Models
ScODA: An Emerging Pipeline for Evaluating Distributed Database Performance To Support Operational Data Analytics
DiffPro: Joint Timestep and Layer-Wise Precision Optimization for Efficient Diffusion Inference
A Quantum Solver for Multidimensional Partial Differential Equations: Practical Case Studies
GPU Kernels for Mixture of Experts
Distributed 3D Gaussian Splatting for High-Resolution Isosurface Visualization
Fast Linear Solvers via AI-Tuned Markov Chain Monte Carlo-Based Matrix Inversion
Real-Time ML-Based Defense Against Malicious Payload in Reconfigurable Embedded Systems
Performance Engineering of Scientific Applications with MVAPICH and TAU Using Emerging Communication Primitives
GATSched: Multi-Objective Graph Attention Networks for Energy-Efficient HPC Job Scheduling
A Scalability Study of Quantum Algorithms for Dimensionality Reduction of Multidimensional Data
Luthier: A Dynamic Binary Instrumentation Framework Targeting AMD GPUs
Template Task-Based Multiresolution Analysis in Hybrid Environments
A Toolbox for Load Balancing Development and Analysis in WarpX/AMReX Applications
An Approach for Correlating Compiler Optimizations with Runtime Performance
Seamless Scaling of Applications Across Programming Models
Orchid: Towards Heterogeneous Batched Eigenvalue Solvers
Bridging the Quantum Coding Gap: Instruction-Tuned LLMs for Qiskit
IncineRate: Multi-Modal FPGA Accelerator for SCNNs
Understanding Communication Bottlenecks in Multi-Node LLM Inference
Learning To Select Scheduling Algorithms in OpenMP
SRAP: Sender-Side Receiver-Aware Port Selection for High-Speed Multi-Flow TCP
TidalMark: A Scalable Benchmark for Coastal Water Level Forecasting
Parallel Local Motif Counting on Large-Scale Dynamic Graphs
Unraveling Distant Galaxies: Analyzing IFU Data with Parsl and Academy
Local vs. Global FFT Approaches for High-Performance Ultrasound Simulation on Multi-GPU Systems
Exploring Fine-Grained Parallelism in Data-Flow Runtime Systems on Many-Core Systems
Evaluating the Usage of Python Libraries on a Production Supercomputer
An Efficient GEMM Acceleration Method for LLM Inference with Variable-Length Sequences
Shortcut Mixup Policy: Toward Improving Robustness and Speed in Goal-Conditioned RL
Understanding LLM Behavior on HPC Data via Mechanistic Interpretability
Optimizing and Extending Periodogram Computations for Astronomy
Scalable Alternative Route Computation with ACE: A C++17 Library for HPC Traffic Simulations
ParaViz3D: MPI Trace Visualization with 3D Video
Range Search on Heterogeneous Systems with Processing-in-Memory Architecture
High-Performance Sparse Attention on Tensor Cores: Fused3S and Beyond
Author
Mojo: Python-Like MLIR-Based GPU Portable Science Kernels
Detecting Silent Data Corruption in Sparse Matrices Using Hardware Performance Counters
Understanding GPU Utilization Using LDMS Data on Perlmutter
Advancing EEG Signal Analysis with Quantum Machine Learning
Harmony: Converged Supercomputer Scratch and Archival Filesystems
AdversaGuard: A Distributed Data Poisoning Benchmark for Parallel AI
High Performance Batch SVD Using GPUs
ChatHPC: Building the Foundations for a Productive and Trustworthy AI-Assisted HPC Ecosystem
Energy-Efficient Multimodal LLM Inference: Stage-Level Characterization and Input-Aware Controls​
Characterizing Performance and Energy Trade-Offs on the Aurora Supercomputer
Hardware-Aware Quantum Circuit Synthesis
Process-Based Predictors of Vulnerability Reintroduction
Distributed Modular Digital Twin Network for High-Performance and Reliable Data Centers
Unified Performance Modeling Stack for Distributed GPU Applications: Complementing Analytical Insights with Machine Learning
CUR-MoE: Portable Mixture-of-Experts with Interpretable High-Ratio Compression
Building the Foundation for Machine Learning-Based Mars Weather Forecasting
Accelerating Scientific Workflows with LLM-Driven Compiler Optimizations for Generated High-Performance Hardware
Memory-Efficient CFD Based on MPS: Effective One-Billion-Cell Resolution on a Single Node
csDF: A Double-Float Arithmetic Library for the Cerebras CS-2
Inference-as-a-Service Prototype at NERSC
Multi-GPU Implementation and Roofline Analysis of a Numerical Global Ocean Model
European Open Web Index: Large Complex Graph Visualization
Mitigating I/O Bottlenecks in LiDAR Pipelines by Directly Merging Neural Decompression and Semantic Segmentation
Numerical Investigation of Radiation Hydrodynamic Instabilities at Scale with FleCSI-HARD
Mixed Compute Environments with OpenCHAMI
Leveraging Large Language Models for Property Prediction in Polymorphic Organic Semiconductors
Scalable Multi-Node Multi-GPU Datalog Engine with Energy-Aware Profiling
Applying Lossy Compression Techniques to GNN Training
CROSS-HPC System Bayesian Optimization with Adaptive Transfer
Configuring Large Language Models for Regional Ocean Model Development
Can Lossy Compression Benefit NVMe-Based I/O?
From Petabytes to Predictions: Harnessing Large-Scale NeuroBlu Mental Health Data and ML To Mitigate Medication Non-Adherence
C++ Standard Parallelism for GPU Programming in a Particle-In-Cell Application
Tensor Core Accelerated Fast Multipole Method for GROMACS
Analyzing Dataset Popularity for Optimizing In-Network Storage
Time-Stepping Hamiltonian Simulation for Solving Nonlinear PDEs via a Quantum-Classical Hybrid Approach
Enabling Real-Time, Extreme-Scale Bayesian Inference: FFT-Based GPU-Accelerated Matrix-Vector Products for Block-Triangular Toeplitz Matrices
Massively Parallel GPU Rasterizer for Next-Generation Computational Lithography
Evaluating the Power-Monitoring Capabilities of Aurora
Divide, Conquer, and Denoise: Hybrid Parallel Diffusion with Memory-Aware Coarse-to-Fine Inference
The Impact of Maximum Vector Length on Cache Management Techniques in RISC-V Vector Extension
Classifying Performance Bounds Using Machine Learning
From Legacy to Portable: An Agentic AI Workflow for Fortran Code Translation and Cross-Architecture Optimization
Wafer-Scale Simulation of Mutator Allele Dynamics in Large Asexual Populations
Job Grouping-Based Intelligent Resource Recommendation Framework
Enhancing Usability and Performance in Experimental Environments Management
Towards Application Agnostic HPC Profiling
Using Hardware Metrics To Understand Performance of the RAJA Performance Suite Kernels in Different GPU Modes on MI300A
A Kokkos-Based Proxy of the Exascale Metagenome Assembler MetaHipMer2: A First Use of Kokkos for Computational Biology