BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201809Z
LOCATION:274
DTSTART;TZID=America/Chicago:20251117T153000
DTEND;TZID=America/Chicago:20251117T155000
UID:submissions.supercomputing.org_SC25_sess215_ws_ai4s115@linklings.com
SUMMARY:FIRST: Federated Inference Resource Scheduling Toolkit for Scienti
 fic AI Model Access
DESCRIPTION:Aditya Tanikanti, Benoit Cote, Yanfei Guo, and Le Chen (Argonn
 e National Laboratory (ANL)); Nickolaus Saint (The University of Chicago);
  Ryan Chard, Ken Raffenetti, Rajeev Thakur, Thomas Uram, and Ian Foster (A
 rgonne National Laboratory (ANL)); Michael E. Papka (Argonne National Labo
 ratory (ANL), University of Illinois Chicago); and Venkatram Vishwanath (A
 rgonne National Laboratory (ANL))\n\nWe present the Federated Inference Re
 source Scheduling Toolkit (FIRST), a framework enabling Inference-as-a-Ser
 vice across distributed High-Performance Computing (HPC) clusters. FIRST p
 rovides cloud-like access to diverse AI models, like Large Language Models
  (LLMs), on existing HPC infrastructure. Leveraging Globus Auth and Globus
  Compute, the system allows researchers to run parallel inference workload
 s via an OpenAI-compliant API on private, secure environments. This cluste
 r-agnostic API allows requests to be distributed across federated clusters
 , targeting numerous hosted models. FIRST supports multiple inference back
 ends (e.g., vLLM), auto-scales resources, maintains "hot" nodes for low-la
 tency execution, and offers both high-throughput batch and interactive mod
 es. The framework addresses the growing demand for private, secure, and sc
 alable AI inference in scientific workflows, allowing researchers to gener
 ate billions of tokens daily on-premises without relying on commercial clo
 ud infrastructure.\n\nRecording: Livestreamed, Recorded\n\nRegistration Ca
 tegory: Technical Program Reg Pass, Workshop Reg Pass\n\nSession Chairs: G
 okcen Kestor (Barcelona Supercomputing Center (BSC); University of Califor
 nia, Merced); Dong Li (University of California, Merced); and Murali Emani
  (Argonne National Laboratory (ANL))\n\n
END:VEVENT
END:VCALENDAR
