Session Full Schedule · Contributors · Organizations · Search Program · My Schedule · Happening NowMore…Search ProgramMy ScheduleHappening NowResearch and ACM SRC Posters: Poster Reception (Research, ACM SRC Grads/Undergrads)Session ChairsKento SatoRIKEN Center for Computational Science (R-CCS)Chris SchlipaliusPawsey Supercomputing Research CentreCommonwealth Scientific and Industrial Research Organisation (CSIRO), AustraliaAnja GerbesGeorg-August-Universität GöttingenEvent TypeResearch and ACM SRC PostersTimeTuesday, 18 November 20255:15pm - 7:00pm CSTLocationSecond Floor AtriumTagsResearch & ACM SRC PostersRegistration Categories TP Similar SessionsBest Poster Finalist Presentations (ACM SRC Grads)Best Poster Finalist Presentations (ACM SRC Undergrads)Doctoral Showcase I PresentationsPresentationsScaling Singular Values Beyond GPU Memory Limits: Out-of-Core, GPU-Accelerated, and Unified Across Data Precision and HardwareAuthorEvelyne RingootEchoes of Earth: Building an Autonomous Environmental Lab for Acoustic SensingAuthorsHudson ReynoldsAlex TueckeMike ShermanKate KeaheyScalable Execution Framework for R on Manycore SystemsAuthorsXiran ZhangJavier ConejeroSameh AbdulahJorge EjarqueYing SunRosa M. BadiaDavid E. KeyesMarc G. GentonSync-Free GPU Parallelization of Sparse Kernels from Sequential Python CodeAuthorMalko-Bani SomoFacilitating Mixed Python-Fortran HPC Codes: 4D Drift-Kinetic Simulations with PyccelAuthorsEmily BourneYaman GüçlüAlgorithms and Applications of Dynamic Network Analysis Using CANDYAuthorsAashish PandeyArindam KhandaS.M. ShovanAli Y. KhanBoyana NorrisSajal K. DasSanjukta BhowmickWhen Label Propagation Outperforms BFS in Breadth-First Graph TraversalAuthorsKalsuda LapborisuthSrinivas AluruA Formal Characterization of Non-Monotonicity in Tensor CoresAuthorsPaul JiangVivian ZhengTowards a GPU-Accelerated Web-Based Graph Rendering Framework for Large-Scale Protein NetworksAuthorsJiaxin LuLandon DykenShilpika ShilpikaVenkatram VishwanathMichael PapkaSidharth KumarUnmasking Performance Variability in GPU Codes on SupercomputersAuthorsCunyang WeiKeshav PradeepAbhinav BhateleAutoSlim: Intelligent Automata Graph Optimization for Efficient AccelerationAuthorsTiffany YuRasha KarakchiShipping HPC Ecosystems Across Platforms: Portable and Composable HPC Clusters as CodeAuthorsGerman Felipe Giraldo VillaThéo GrivelGeorge IoannidisEdita KizinevicCarolina LindqvistNicolas LitchinkoPablo LlopisAntonio Javier RussoGilles FouresteyChameleon Concierge: Retrieval-Augmented Generation (RAG) To Enhance Open Testbed DocumentationAuthorSaieda Ali ZadaExplicit Low-Order Finite-Element Wave Simulation Accelerated with Variable-Precision Computing Using INT8 Tensor CoresAuthorsKohei FujitaTsuyoshi IchimuraMuneo HoriLalith MaddegedaraVaultX Merge: Breaking Memory Barriers in Proof-of-Space Plot GenerationAuthorsArnav SirigereVarvara BondarenkoIoan RaicuOptimizing Collectives with Large Payloads on GPU-Based SupercomputersAuthorsSiddharth SinghMahua SinghKeshav PradeepAbhinav BhateleWiCAT: Reducing Congestion at Wireless Interfaces in Heterogeneous ArchitecturesAuthorTarun SharmaIntelligent Surrogates Pay Attention to Data, Improving Multi-Objective HPC OptimizationAuthorsAshna Nawar AhmedBanooqa BandayTerry JonesTanzima Z. IslamCIRE: LLVM Analysis for Floating-Point Rounding Error Affected by Precision and OptimizationsAuthorsCayden LundTanmay TirpankarGanesh GopalakrishnanGNNs on Evolving Graphs: A Benchmark of Incremental Updates and Meta-Learning ApproachesAuthorsSriram SrinivasanSanjukta BhowmickHamdan AlabsiRand ObeidatAn Agent-Based Viral Venture: Adaptive Tool Selection for Scalable GenomicsAuthorsNaomi KolodisnerAlok Kamatar (Advisor)J. Greg Pauloski (Advisor)PhySiViT: A Physics Simulation Vision TransformerAuthorsJessica EzembaJames AffulMei-Yu WangEnabling Efficient Runtime Data Analysis to a Crystal Deformation SimulationAuthorArthur JaquardCompute System Simulator: Modeling the Impact of Allocation Policy and Hardware Reliability on HPC Cloud Resource UtilizationAuthorsJarrod LeddyHuseyin YildizDivergence Prediction System for CFD SimulationsAuthorsTakashi SogaTakanori UchidaSusumu DateJACC: Easy CPU/GPU Performance Portability for Scientific Applications in JuliaAuthorsWilliam GodoyPedro Valero-LaraPhilip FacklerKeita TeranishiJeffrey VetterJhonny GonzalezJose GonzalezAlexis HuanteMassively Parallel Bayesian Inference Framework for GPU Supercomputers: Application to Estimation of Coseismic Fault SlipAuthorsKai NakaoTsuyoshi IchimuraKohei FujitaOptimizing Task-Driven Offloading in LLVMAuthorsJan KrausJoachim JenkeChristian TerbovenHydraCache: LLM Inference Prefill Parallelization Through Distributed Cache BlendingAuthorsAdib Rezaei ShahmirzadiShayan ShabihiMona MoghadampanahFurong HuangDimitrios S. NikolopoulosMPI-SGX: Enabling Confidential Computing for MPI Parallel Applications with Intel SGX TechnologyAuthorsKota ShimojimaHayato YamakiHiroki HondaShinichiro MatsuoAtsuko TakefusaShinobu MiwaHeterogeneity-Aware Task Allocation for Modern HPC SystemsAuthorsSowmya YellapragadaJessica Imlau DagostiniKevin GottRebecca Hartman-BakerProductive Scalable Distributed Task Scheduling Using an MPI-based Backend for DaggerAuthorYan GuimarãesOptimizing the GPU All-Reduce Using Multiple Processes Per GPUAuthorsMichael AdamsAmanda BienzDiOMP-Offloading: Portable OpenMP Offloading for Distributed Heterogeneous SystemsAuthorsBaodi ShanMauricio Araya-PoloBarbara ChapmanCATIOS: Time-Resolved I/O-Aware Job Scheduling for HPC SystemsAuthorsYuTsen TsengMasatoshi KawaiKeichi TakahashiHiroyuki TakizawaCan Long-Haul RDMA Benefit Federated Learning?AuthorsZhonghao ChenYuke LiDuo ZhangXiaoyi LuAccelerating AI Co-Scientists with HPC InfrastructureAuthorSuryatejas AppanaForward Error Bounds and Efficient Algorithms for Computing a Tensor Times Matrix Chain in Low Precision on GPUsAuthorsJulian BellavitaPiyush SaoRamakrishnan KannanBetween the NIC and a Hard Place: Evaluating 400 Gb/s Ethernet for HPC Data TransfersAuthorsAdelle FerrisEvelyn NeedhamNikole GrandezJesse MartinezDoug EganJulia with Intelligent Runtime for Heterogeneous ComputingAuthorsNarasinga Rao MiniskarPedro Valero-LaraWilliam GodoyKeita TeranishiJeffrey S. VetterWONDERS: Integrating WOW, PONDER, and SCALE for Enhanced Scheduling PerformanceAuthorsFabian LehmannJonathan RauJonathan BaderOdej KaoUlf LeserEvaluating LiDAR Compression for 3D Semantic Segmentation in Diverse Off-Road Environments on GOOSE DatasetAuthorsAdam NiemczuraMax FaykusOyinlolu OdetoyeMelissa SmithJon CalhounScott GroelNovel Graph Alignment Algorithms for Identifying Non-Determinism in Large-Scale SimulationsAuthorDhroov PandeyAccelerating Linear Solve with Mixed Precision Nested Recursive Subdivision on AI HardwareAuthorVicki CarricaPractical Viability of Translating Legacy Fortran Code to C++ Using Large Language ModelsAuthorsRen ImaiMasatoshi KawaiKeichi TakahashiHiroyuki TakizawaScODA: An Emerging Pipeline for Evaluating Distributed Database Performance To Support Operational Data AnalyticsAuthorsNicholas SynovicFNU ShilpikaSilvio RizziDoug WaldronGeorge K. ThiruvathukalMichael E. PapkaDiffPro: Joint Timestep and Layer-Wise Precision Optimization for Efficient Diffusion InferenceAuthorsFarhana AminKanchon GharamiDimitrios S. NikolopoulosA Quantum Solver for Multidimensional Partial Differential Equations: Practical Case StudiesAuthorsManu ChaudharyKareem El-ArabyAlvir NobelIshraq IslamManish SinghSunday OgundeleKieran EganSneha ThomasVincent VordtriedeDevon BontragerSerom KimEsam El-ArabyGPU Kernels for Mixture of ExpertsAuthorsArthur FeeneyYing Wai LiAparna ChandramowlishwaranDistributed 3D Gaussian Splatting for High-Resolution Isosurface VisualizationAuthorsMengjiao HanAndres SewellJoseph InsleyJanet KnowlesVictor A. MateevitsiMichael E. PapkaSteve PetruzzaSilvio RizziFast Linear Solvers via AI-Tuned Markov Chain Monte Carlo-Based Matrix InversionAuthorsAnton LebedevWon Kyung LeeSoumyadip GhoshOlha I. YamanVassilis KalantzisYingdong LuTomasz NowickiShashanka UbaruLior HoreshVassil AlexandrovReal-Time ML-Based Defense Against Malicious Payload in Reconfigurable Embedded SystemsAuthorsRye Stahle-SmithRasha KarakchiPerformance Engineering of Scientific Applications with MVAPICH and TAU Using Emerging Communication PrimitivesAuthorsDhabaleswar K. (DK) PandaSameer ShendeAhmad AbdelfattahYifeng CuiGATSched: Multi-Objective Graph Attention Networks for Energy-Efficient HPC Job SchedulingAuthorKyrian AdimoraA Scalability Study of Quantum Algorithms for Dimensionality Reduction of Multidimensional DataAuthorsKareem El-ArabyThom PopovicAlvir NobelSunday OgundeleKatherine KlymkoDaan CampsAnastasiia ButkoEsam El-ArabyLuthier: A Dynamic Binary Instrumentation Framework Targeting AMD GPUsAuthorsMatin Raayai-ArdakaniNorman RubinDavid KaeliTemplate Task-Based Multiresolution Analysis in Hybrid EnvironmentsAuthorsNilesh ChaturvediJoseph SchuchartRobert J. HarrisonA Toolbox for Load Balancing Development and Analysis in WarpX/AMReX ApplicationsAuthorsJessica Imlau DagostiniSowmya YellapragadaKevin GottRebecca Hartman-BakerAn Approach for Correlating Compiler Optimizations with Runtime PerformanceAuthorsBefikir BogaleOlga PearceTom ScoglandMichela TauferSeamless Scaling of Applications Across Programming ModelsAuthorsReto KrummenacherQuentin GuilloteauJonas H. Müller KorndörferFlorina M. CiorbaOrchid: Towards Heterogeneous Batched Eigenvalue SolversAuthorMatthew ChungBridging the Quantum Coding Gap: Instruction-Tuned LLMs for QiskitAuthorsSixu ChenYuqi ZhangQiang GuanIncineRate: Multi-Modal FPGA Accelerator for SCNNsAuthorsBjörn A. LindqvistArtur PodobasUnderstanding Communication Bottlenecks in Multi-Node LLM InferenceAuthorsPrajwal SinghaniaSiddharth SinghLannie Dalton HoughIshan RevankarHarshitha MenonCharles JekelAbhinav BhateleLearning To Select Scheduling Algorithms in OpenMPAuthorsJonas H. Müller KorndörferAli MohammedAhmed EleliemyQuentin GuilloteauReto KrummenacherFlorina CiorbaSRAP: Sender-Side Receiver-Aware Port Selection for High-Speed Multi-Flow TCPAuthorsShingo HattoriOsamu TatebeTidalMark: A Scalable Benchmark for Coastal Water Level ForecastingAuthorsLucas RaicuDaniel GrzendaIan FosterKyle ChardParallel Local Motif Counting on Large-Scale Dynamic GraphsAuthorsAli KhanSanjukta BhowmickMichela TauferUnraveling Distant Galaxies: Analyzing IFU Data with Parsl and AcademyAuthorDaniel BabniggLocal vs. Global FFT Approaches for High-Performance Ultrasound Simulation on Multi-GPU SystemsAuthorsOliver KuníkJiri JarosExploring Fine-Grained Parallelism in Data-Flow Runtime Systems on Many-Core SystemsAuthorsWenyi WangMaxime GonthierHaibin LaiPoornima NookalaHaochen PanIan FosterIoan RaicuKyle ChardEvaluating the Usage of Python Libraries on a Production SupercomputerAuthorThomas PapkaAn Efficient GEMM Acceleration Method for LLM Inference with Variable-Length SequencesAuthorsYu ZhangLu LuShortcut Mixup Policy: Toward Improving Robustness and Speed in Goal-Conditioned RLAuthorsMatthew HyattYassir AtlasHal BryntesonDiego Roa PerdomoAthena AngaraMengjiao HanJoseph InsleyJanet KnowlesYongho KimVictor MateevitsiMichael PapkaSilvio RizziGeorge ThiruvathukalNicola FerrierUnderstanding LLM Behavior on HPC Data via Mechanistic InterpretabilityAuthorsMd Mahbubur RahmanArjun GuhaHarshitha MenonOptimizing and Extending Periodogram Computations for AstronomyAuthorsYuwei SunLehman GarrisonScalable Alternative Route Computation with ACE: A C++17 Library for HPC Traffic SimulationsAuthorsPaulo SilvaPavlína SmolkováKateřina SlaninováJan MartinovičJoão BarbosaMatej ŠpeťkoEmanuele VitaliParaViz3D: MPI Trace Visualization with 3D VideoAuthorsJean-Yves VerhaegheGeorg HagerAyesha AfzalRange Search on Heterogeneous Systems with Processing-in-Memory ArchitectureAuthorsTasmia JannatSatish PuriMichael GowanlockHigh-Performance Sparse Attention on Tensor Cores: Fused3S and BeyondAuthorZitong LiMojo: Python-Like MLIR-Based GPU Portable Science KernelsAuthorTatiana MelnichenkoDetecting Silent Data Corruption in Sparse Matrices Using Hardware Performance CountersAuthorsMinseop ChoiOrlando AriasSeung Woo SonUnderstanding GPU Utilization Using LDMS Data on PerlmutterAuthorsOnur CankurBrian AustinAbhinav BhateleAdvancing EEG Signal Analysis with Quantum Machine LearningAuthorsStephanie MurrayErika ParsonsHarmony: Converged Supercomputer Scratch and Archival FilesystemsAuthorJake CarrollAdversaGuard: A Distributed Data Poisoning Benchmark for Parallel AIAuthorsYulia KumarSolomon ThomasDejaun GayleJ. Jenny LiDov KrugerHigh Performance Batch SVD Using GPUsAuthorAhmad AbdelfattahChatHPC: Building the Foundations for a Productive and Trustworthy AI-Assisted HPC EcosystemAuthorsPedro Valero-LaraAaron YoungMohammad Alaul Haque MonilSwaroop PophaleZheming JinJeffrey S. VetterKeita TeranishiWilliam F. GodoyEnergy-Efficient Multimodal LLM Inference: Stage-Level Characterization and Input-Aware ControlsAuthorsMona MoghadampanahAdib Rezaei ShahmirzadiDimitrios S. NikolopoulosCharacterizing Performance and Energy Trade-Offs on the Aurora SupercomputerAuthorsSolomon BekeleSwann PerarnauBrice VideauHardware-Aware Quantum Circuit SynthesisAuthorsNathan JonesAkhilesh BondapalliToby CoxIan LewisRong GeProcess-Based Predictors of Vulnerability ReintroductionAuthorsSamiha ShimmiNicholas SynovicMona RahimiGeorge ThiruvathukalDistributed Modular Digital Twin Network for High-Performance and Reliable Data CentersAuthorsYan ChenXing LuCary FaulknerAlex VlachokostasHanlong WanJeremy LerondUnified Performance Modeling Stack for Distributed GPU Applications: Complementing Analytical Insights with Machine LearningAuthorUrvij SaroliyaCUR-MoE: Portable Mixture-of-Experts with Interpretable High-Ratio CompressionAuthorRitesh BhirudBuilding the Foundation for Machine Learning-Based Mars Weather ForecastingAuthorMohammad AltiwainyAccelerating Scientific Workflows with LLM-Driven Compiler Optimizations for Generated High-Performance HardwareAuthorsRobert RamstadNicolas Bohm AgostiniAntonino TumeoMemory-Efficient CFD Based on MPS: Effective One-Billion-Cell Resolution on a Single NodeAuthorsJunya OnishiAyato TakiiSangwon KimYounghwa ChoMakoto TsubokuracsDF: A Double-Float Arithmetic Library for the Cerebras CS-2AuthorsReo NagashimaAkeru NakamuraKai MurakamiRyunosuke MatsuzakiDaichi MukunokiTakaaki MiyajimaInference-as-a-Service Prototype at NERSCAuthorsColin ThomasPo-Han HuangHilary UtaegbulamJohannes BlaschkeBruno CoimbraPengfei DingXiangyang JuAndrew NaylorMichael WangMulti-GPU Implementation and Roofline Analysis of a Numerical Global Ocean ModelAuthorsTakateru YamagishiMasao KurogiTakao KawasakiYoshimasa MatsumuraHiroyasu HasumiEuropean Open Web Index: Large Complex Graph VisualizationAuthorsPavlina SmolkovaKaterina SlaninovaMitigating I/O Bottlenecks in LiDAR Pipelines by Directly Merging Neural Decompression and Semantic SegmentationAuthorsEthan MarquezMax FaykusOyinlolu OdetoyeMelissa SmithJon CalhounNumerical Investigation of Radiation Hydrodynamic Instabilities at Scale with FleCSI-HARDAuthorsMåns I. AnderssonIsaac C. BannermanMoon B. HazarikaAkshit JariwalaJonathan MathurinMadela B. QuashieJulien LoiseauHyun LimMixed Compute Environments with OpenCHAMIAuthorsSean GibsonRichard KimSamuel QuanTravis CottonThomas MackellLeveraging Large Language Models for Property Prediction in Polymorphic Organic SemiconductorsAuthorsShreya PagariaMei-Yu WangDana O’ConnorJulian UranPaola BuitragoScalable Multi-Node Multi-GPU Datalog Engine with Energy-Aware ProfilingAuthorsAhmedur Rahman ShovonSidharth KumarApplying Lossy Compression Techniques to GNN TrainingAuthorsMilan ShahReece NeffMichela BecchiCROSS-HPC System Bayesian Optimization with Adaptive TransferAuthorsAbrar HossainKishwar AhmedConfiguring Large Language Models for Regional Ocean Model DevelopmentAuthorsAidan JanneyGiovanni Seijo-EllisDan AmrheinCan Lossy Compression Benefit NVMe-Based I/O?AuthorsDarren NgDuo ZhangSheng DiZhaorui ZhangXiaoyi LuFrom Petabytes to Predictions: Harnessing Large-Scale NeuroBlu Mental Health Data and ML To Mitigate Medication Non-AdherenceAuthorsAlyson CollinsCathy SandovalMaya SeshanSrishti SrivastavaJosh McWilliamsC++ Standard Parallelism for GPU Programming in a Particle-In-Cell ApplicationAuthorsEster El KhouryMathieu LobetJulien BigotLaurent ColombetTensor Core Accelerated Fast Multipole Method for GROMACSAuthorsJiamian HuangMuhammad Umair SadiqRio YokotaBerk HessAnalyzing Dataset Popularity for Optimizing In-Network StorageAuthorsGunwoo KimAlex SimKesheng WuTime-Stepping Hamiltonian Simulation for Solving Nonlinear PDEs via a Quantum-Classical Hybrid ApproachAuthorsSangwon KimJunya OnishiAyato TakiiYounghwa ChoTsubokura MakotoEnabling Real-Time, Extreme-Scale Bayesian Inference: FFT-Based GPU-Accelerated Matrix-Vector Products for Block-Triangular Toeplitz MatricesAuthorsSreeram VenkatOmar GhattasMassively Parallel GPU Rasterizer for Next-Generation Computational LithographyAuthorsLoay HegazyMohamed TaherSherif HammoudaEvaluating the Power-Monitoring Capabilities of AuroraAuthorPrecious EyabiDivide, Conquer, and Denoise: Hybrid Parallel Diffusion with Memory-Aware Coarse-to-Fine InferenceAuthorsFarhana AminKanchon GharamiDimitrios NikolopoulosThe Impact of Maximum Vector Length on Cache Management Techniques in RISC-V Vector ExtensionAuthorsShunya NomuraJiaheng LiuKeichi TakahashiHiroyuki TakizawaClassifying Performance Bounds Using Machine LearningAuthorsLewis LittmanTom DeakinFrom Legacy to Portable: An Agentic AI Workflow for Fortran Code Translation and Cross-Architecture OptimizationAuthorsSparsh GuptaKamalavasan KamalakkannanMaxim MoraruGalen ShipmanPatrick DiehlWafer-Scale Simulation of Mutator Allele Dynamics in Large Asexual PopulationsAuthorsMatthew Andres MorenoEmily DolsonLuis ZamanJob Grouping-Based Intelligent Resource Recommendation FrameworkAuthorsBeste OztopBenjamin SchwallerVitus J. LeungJim BrandtBrian KulisManuel EgeleAyse K. CoskunEnhancing Usability and Performance in Experimental Environments ManagementAuthorsZahra TemoriPaul MarshallKate KeaheyTowards Application Agnostic HPC ProfilingAuthorsHari Teja JajulaDhruva KulkarniBrian AustinPurushotham BangaloreUsing Hardware Metrics To Understand Performance of the RAJA Performance Suite Kernels in Different GPU Modes on MI300AAuthorsAmr AbouelmagdStephanie BrinkMichael McKinseyDavid BoehmeJason BurmarkBrian RyujinTom ScoglandOlga PearceA Kokkos-Based Proxy of the Exascale Metagenome Assembler MetaHipMer2: A First Use of Kokkos for Computational BiologyAuthorsLogan WilliamsGavin ConantMichela BecchiJan CieskoAmy Powell