BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201258Z
LOCATION:Second Floor Atrium
DTSTART;TZID=America/Chicago:20251121T080000
DTEND;TZID=America/Chicago:20251121T120000
UID:submissions.supercomputing.org_SC25_sess620_post304@linklings.com
SUMMARY:HydraCache: LLM Inference Prefill Parallelization Through Distribu
 ted Cache Blending
DESCRIPTION:Adib Rezaei Shahmirzadi (Virginia Tech), Shayan Shabihi (Unive
 rsity of Maryland), Mona Moghadampanah (Virginia Tech), Furong Huang (Univ
 ersity of Maryland), and Dimitrios S. Nikolopoulos (Virginia Tech)\n\nThe 
 prefill phase of large language model (LLM) inference, where the input pro
 mpt is processed to generate a key-value (KV) cache, is a critical latency
  bottleneck for input sequences. Existing serving architectures face a tra
 de-off: data parallelism (DP) offers flexibility but cannot accelerate a s
 ingle long prompt, while tensor parallelism (TP) parallelizes prefill but 
 at the cost of rigid resource allocation and constant communication overhe
 ad at each layer. We introduce HydraCache, a system that resolves this pro
 blem by enabling a cluster of independent, data-parallel model replicas to
  collaborate on-demand to parallelize the prefill of a single long prompt.
  Our core contribution is DistBlendAttention, a lightweight mechanism that
  fuses distributed KV caches with minimal communication, avoiding the proh
 ibitive overheads of both TP and traditional sequence parallelism. Our eva
 luation shows that HydraCache significantly reduces Time-to-First-Token (T
 TFT) up to 7x for requests and enables flexible, SLO-aware serving.\n\nTag
 : Research & ACM SRC Posters\n\nRegistration Category: Technical Program R
 eg Pass\n\nSession Chairs: Kento Sato (RIKEN Center for Computational Scie
 nce (R-CCS)); Anja Gerbes (Georg-August-Universität Göttingen); and Chris 
 Schlipalius (Pawsey Supercomputing Research Centre; Commonwealth Scientifi
 c and Industrial Research Organisation (CSIRO), Australia)\n\n
END:VEVENT
END:VCALENDAR
