BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201805Z
LOCATION:230
DTSTART;TZID=America/Chicago:20251120T114500
DTEND;TZID=America/Chicago:20251120T120000
UID:submissions.supercomputing.org_SC25_sess534_drs116@linklings.com
SUMMARY:Advancing Data Center Workloads with Data Processing Units
DESCRIPTION:Arjun Kashyap (University of California, Merced)\n\nModern dat
 a center workloads demand substantial server resources, motivating the ado
 ption of data processing units (DPUs) for improved efficiency. Despite inc
 reasing deployment, systematic characterization of SoC-based DPUs remains 
 limited. We present a rigorous evaluation of NVIDIA’s BlueField-1, BlueFie
 ld-2 (BF-2), and BlueField-3 (BF-3) across 15 benchmarks, revealing key id
 iosyncrasies in network, DMA, and memory. We further provide design recomm
 endations and release our artifacts to the community. Additionally, naivel
 y integrating DPUs into workloads often reduces server resource usage with
 out necessarily delivering high performance. In-memory key-value stores (K
 VS) are widely used for edge data storage, where low latency and high thro
 ughput are essential. We explore fine-grained offloading of in-memory CPU-
 based KVS to SoC-based DPUs by decomposing KVS and offloading the communic
 ation engine, the most CPU-intensive component, to enhance performance. We
  also propose a series of performance optimizations, such as overlapped re
 quest/response handling, reduced DMA operations, and dual communication en
 gines. Our design achieves up to 68% lower latency and 36% higher throughp
 ut compared to CPU-only or coarse-grained offloading.\n\nApplications in c
 ontainers or VMs commonly rely on TCP/IP for communication in HPC clouds a
 nd data centers, yet TCP/IP introduces significant bottlenecks for NVMe-ov
 er-Fabrics I/O in disaggregated storage. We propose NVMe-over-Adaptive-Fab
 ric (NVMe-oAF), an adaptive communication channel that leverages locality 
 awareness and optimized shared memory/TCP paths to accelerate I/O-intensiv
 e workloads. Co-designed with Intel’s SPDK, NVMe-oAF achieves up to 7.1x h
 igher bandwidth and 4.2x lower latency compared to TCP/IP over commodity E
 thernet (10–100 Gbps), while delivering up to 7x bandwidth gains for HDF5 
 applications when integrated with H5bench.\n\nTag: Research & ACM SRC Post
 ers\n\nRecording: Livestreamed, Recorded\n\nRegistration Category: Technic
 al Program Reg Pass\n\nSession Chairs: Yanfei Guo (Argonne National Labora
 tory (ANL)); Shirley Moore (University of Texas at El Paso); Kento Sato (R
 IKEN Center for Computational Science (R-CCS)); Chris Schlipalius (Pawsey 
 Supercomputing Research Centre; Commonwealth Scientific and Industrial Res
 earch Organisation (CSIRO), Australia); and Anja Gerbes (Georg-August-Univ
 ersität Göttingen)\n\n
END:VEVENT
END:VCALENDAR
