Close

Session

Event Type
Research and ACM SRC Posters
TimeFriday, 21 November 20258:00am - 12:00pm CST
LocationSecond Floor Atrium
Tags
Research & ACM SRC Posters
Registration Categories
TP
Presentations
Intelligent Surrogates Pay Attention to Data, Improving Multi-Objective HPC Optimization
Numerical Investigation of Radiation Hydrodynamic Instabilities at Scale with FleCSI-HARD
High-Performance Sparse Attention on Tensor Cores: Fused3S and Beyond
Author
A Quantum Solver for Multidimensional Partial Differential Equations: Practical Case Studies
Evaluating the Power-Monitoring Capabilities of Aurora
DiOMP-Offloading: Portable OpenMP Offloading for Distributed Heterogeneous Systems
Performance Engineering of Scientific Applications with MVAPICH and TAU Using Emerging Communication Primitives
Enhancing Usability and Performance in Experimental Environments Management
Can Lossy Compression Benefit NVMe-Based I/O?
WONDERS: Integrating WOW, PONDER, and SCALE for Enhanced Scheduling Performance
Tensor Core Accelerated Fast Multipole Method for GROMACS
Echoes of Earth: Building an Autonomous Environmental Lab for Acoustic Sensing
Mitigating I/O Bottlenecks in LiDAR Pipelines by Directly Merging Neural Decompression and Semantic Segmentation
Accelerating Scientific Workflows with LLM-Driven Compiler Optimizations for Generated High-Performance Hardware
AutoSlim: Intelligent Automata Graph Optimization for Efficient Acceleration
Job Grouping-Based Intelligent Resource Recommendation Framework
Parallel Local Motif Counting on Large-Scale Dynamic Graphs
DiffPro: Joint Timestep and Layer-Wise Precision Optimization for Efficient Diffusion Inference
ScODA: An Emerging Pipeline for Evaluating Distributed Database Performance To Support Operational Data Analytics
Multi-GPU Implementation and Roofline Analysis of a Numerical Global Ocean Model
Harmony: Converged Supercomputer Scratch and Archival Filesystems
Practical Viability of Translating Legacy Fortran Code to C++ Using Large Language Models
Scaling Singular Values Beyond GPU Memory Limits: Out-of-Core, GPU-Accelerated, and Unified Across Data Precision and Hardware
Julia with Intelligent Runtime for Heterogeneous Computing
Forward Error Bounds and Efficient Algorithms for Computing a Tensor Times Matrix Chain in Low Precision on GPUs
Novel Graph Alignment Algorithms for Identifying Non-Determinism in Large-Scale Simulations
HydraCache: LLM Inference Prefill Parallelization Through Distributed Cache Blending
Energy-Efficient Multimodal LLM Inference: Stage-Level Characterization and Input-Aware Controls​
Local vs. Global FFT Approaches for High-Performance Ultrasound Simulation on Multi-GPU Systems
CUR-MoE: Portable Mixture-of-Experts with Interpretable High-Ratio Compression
Towards Application Agnostic HPC Profiling
Compute System Simulator: Modeling the Impact of Allocation Policy and Hardware Reliability on HPC Cloud Resource Utilization
Mixed Compute Environments with OpenCHAMI
PhySiViT: A Physics Simulation Vision Transformer
Template Task-Based Multiresolution Analysis in Hybrid Environments
Scalable Alternative Route Computation with ACE: A C++17 Library for HPC Traffic Simulations
Mojo: Python-Like MLIR-Based GPU Portable Science Kernels
Real-Time ML-Based Defense Against Malicious Payload in Reconfigurable Embedded Systems
Bridging the Quantum Coding Gap: Instruction-Tuned LLMs for Qiskit
Shipping HPC Ecosystems Across Platforms: Portable and Composable HPC Clusters as Code
Configuring Large Language Models for Regional Ocean Model Development
Using Hardware Metrics To Understand Performance of the RAJA Performance Suite Kernels in Different GPU Modes on MI300A
Memory-Efficient CFD Based on MPS: Effective One-Billion-Cell Resolution on a Single Node
Enabling Real-Time, Extreme-Scale Bayesian Inference: FFT-Based GPU-Accelerated Matrix-Vector Products for Block-Triangular Toeplitz Matrices
A Formal Characterization of Non-Monotonicity in Tensor Cores
Sync-Free GPU Parallelization of Sparse Kernels from Sequential Python Code
Heterogeneity-Aware Task Allocation for Modern HPC Systems
Building the Foundation for Machine Learning-Based Mars Weather Forecasting
Accelerating AI Co-Scientists with HPC Infrastructure
Divergence Prediction System for CFD Simulations
Optimizing the GPU All-Reduce Using Multiple Processes Per GPU
Process-Based Predictors of Vulnerability Reintroduction
Shortcut Mixup Policy: Toward Improving Robustness and Speed in Goal-Conditioned RL
Orchid: Towards Heterogeneous Batched Eigenvalue Solvers
Applying Lossy Compression Techniques to GNN Training
Characterizing Performance and Energy Trade-Offs on the Aurora Supercomputer
Learning To Select Scheduling Algorithms in OpenMP
Massively Parallel Bayesian Inference Framework for GPU Supercomputers: Application to Estimation of Coseismic Fault Slip
Optimizing Task-Driven Offloading in LLVM
Luthier: A Dynamic Binary Instrumentation Framework Targeting AMD GPUs
TidalMark: A Scalable Benchmark for Coastal Water Level Forecasting
Seamless Scaling of Applications Across Programming Models
Wafer-Scale Simulation of Mutator Allele Dynamics in Large Asexual Populations
Distributed Modular Digital Twin Network for High-Performance and Reliable Data Centers
Exploring Fine-Grained Parallelism in Data-Flow Runtime Systems on Many-Core Systems
AdversaGuard: A Distributed Data Poisoning Benchmark for Parallel AI
The Impact of Maximum Vector Length on Cache Management Techniques in RISC-V Vector Extension
CROSS-HPC System Bayesian Optimization with Adaptive Transfer
Facilitating Mixed Python-Fortran HPC Codes: 4D Drift-Kinetic Simulations with Pyccel
When Label Propagation Outperforms BFS in Breadth-First Graph Traversal
Analyzing Dataset Popularity for Optimizing In-Network Storage
Massively Parallel GPU Rasterizer for Next-Generation Computational Lithography
Explicit Low-Order Finite-Element Wave Simulation Accelerated with Variable-Precision Computing Using INT8 Tensor Cores
WiCAT: Reducing Congestion at Wireless Interfaces in Heterogeneous Architectures
MPI-SGX: Enabling Confidential Computing for MPI Parallel Applications with Intel SGX Technology
VaultX Merge: Breaking Memory Barriers in Proof-of-Space Plot Generation
An Approach for Correlating Compiler Optimizations with Runtime Performance
IncineRate: Multi-Modal FPGA Accelerator for SCNNs
Optimizing and Extending Periodogram Computations for Astronomy
GNNs on Evolving Graphs: A Benchmark of Incremental Updates and Meta-Learning Approaches
Detecting Silent Data Corruption in Sparse Matrices Using Hardware Performance Counters
Inference-as-a-Service Prototype at NERSC
From Legacy to Portable: An Agentic AI Workflow for Fortran Code Translation and Cross-Architecture Optimization
Unmasking Performance Variability in GPU Codes on Supercomputers
Leveraging Large Language Models for Property Prediction in Polymorphic Organic Semiconductors
Scalable Execution Framework for R on Manycore Systems
C++ Standard Parallelism for GPU Programming in a Particle-In-Cell Application
Between the NIC and a Hard Place: Evaluating 400 Gb/s Ethernet for HPC Data Transfers
High Performance Batch SVD Using GPUs
Distributed 3D Gaussian Splatting for High-Resolution Isosurface Visualization
GPU Kernels for Mixture of Experts
Divide, Conquer, and Denoise: Hybrid Parallel Diffusion with Memory-Aware Coarse-to-Fine Inference
CATIOS: Time-Resolved I/O-Aware Job Scheduling for HPC Systems
Unraveling Distant Galaxies: Analyzing IFU Data with Parsl and Academy
Understanding GPU Utilization Using LDMS Data on Perlmutter
A Scalability Study of Quantum Algorithms for Dimensionality Reduction of Multidimensional Data
Unified Performance Modeling Stack for Distributed GPU Applications: Complementing Analytical Insights with Machine Learning
Fast Linear Solvers via AI-Tuned Markov Chain Monte Carlo-Based Matrix Inversion
Hardware-Aware Quantum Circuit Synthesis
ParaViz3D: MPI Trace Visualization with 3D Video
Chameleon Concierge: Retrieval-Augmented Generation (RAG) To Enhance Open Testbed Documentation
Advancing EEG Signal Analysis with Quantum Machine Learning
Towards a GPU-Accelerated Web-Based Graph Rendering Framework for Large-Scale Protein Networks
From Petabytes to Predictions: Harnessing Large-Scale NeuroBlu Mental Health Data and ML To Mitigate Medication Non-Adherence
Optimizing Collectives with Large Payloads on GPU-Based Supercomputers
A Kokkos-Based Proxy of the Exascale Metagenome Assembler MetaHipMer2: A First Use of Kokkos for Computational Biology
Understanding Communication Bottlenecks in Multi-Node LLM Inference
European Open Web Index: Large Complex Graph Visualization
Evaluating the Usage of Python Libraries on a Production Supercomputer
CIRE: LLVM Analysis for Floating-Point Rounding Error Affected by Precision and Optimizations
csDF: A Double-Float Arithmetic Library for the Cerebras CS-2
SRAP: Sender-Side Receiver-Aware Port Selection for High-Speed Multi-Flow TCP
Can Long-Haul RDMA Benefit Federated Learning?
Algorithms and Applications of Dynamic Network Analysis Using CANDY
Understanding LLM Behavior on HPC Data via Mechanistic Interpretability
ChatHPC: Building the Foundations for a Productive and Trustworthy AI-Assisted HPC Ecosystem
Productive Scalable Distributed Task Scheduling Using an MPI-based Backend for Dagger
Evaluating LiDAR Compression for 3D Semantic Segmentation in Diverse Off-Road Environments on GOOSE Dataset
Time-Stepping Hamiltonian Simulation for Solving Nonlinear PDEs via a Quantum-Classical Hybrid Approach
An Agent-Based Viral Venture: Adaptive Tool Selection for Scalable Genomics
Accelerating Linear Solve with Mixed Precision Nested Recursive Subdivision on AI Hardware
Scalable Multi-Node Multi-GPU Datalog Engine with Energy-Aware Profiling
Enabling Efficient Runtime Data Analysis to a Crystal Deformation Simulation
A Toolbox for Load Balancing Development and Analysis in WarpX/AMReX Applications
An Efficient GEMM Acceleration Method for LLM Inference with Variable-Length Sequences
Classifying Performance Bounds Using Machine Learning
GATSched: Multi-Objective Graph Attention Networks for Energy-Efficient HPC Job Scheduling
Range Search on Heterogeneous Systems with Processing-in-Memory Architecture
JACC: Easy CPU/GPU Performance Portability for Scientific Applications in Julia