BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201248Z
LOCATION:Second Floor Atrium
DTSTART;TZID=America/Chicago:20251120T080000
DTEND;TZID=America/Chicago:20251120T170000
UID:submissions.supercomputing.org_SC25_sess533_post266@linklings.com
SUMMARY:Shortcut Mixup Policy: Toward Improving Robustness and Speed in Go
 al-Conditioned RL
DESCRIPTION:Matthew Hyatt (Loyola University Chicago, Argonne National Lab
 oratory (ANL)); Yassir Atlas, Hal Brynteson, Diego Roa Perdomo, Athena Ang
 ara, Mengjiao Han, Joseph Insley, Janet Knowles, Yongho Kim, Victor Mateev
 itsi, Michael Papka, and Silvio Rizzi (Argonne National Laboratory (ANL));
  George Thiruvathukal (Loyola University Chicago, Argonne National Laborat
 ory (ANL)); and Nicola Ferrier (Argonne National Laboratory (ANL))\n\nNeur
 al networks trained on large datasets can be effective policies for the co
 ntrol of robotic manipulators. Using self-supervised learning, these netwo
 rks can achieve near-perfect success rates on complex pick-and-place-style
  tasks. However, the speed of task completion is often a barrier to making
  learned policies practical for deployment. For instance, tasks that requi
 re 500 distinct token predictions will require many forward passes through
  the network, in real time. Moreover, to learn optimal task behavior—as in
  reinforcement learning—would require state value assignment across a long
  time horizon. This is often an impediment to learning. To address these c
 hallenges, we present Shortcut Mixup Policy, a method to artificially redu
 ce the task horizon length. Our method consists of training a model on nex
 t-token prediction tasks optionally conditioned on a target state-shortcut
  size. We present initial results using Shortcut Mixup Policy and propose 
 future directions for improvement.\n\nTag: Research & ACM SRC Posters\n\nR
 egistration Category: Technical Program Reg Pass\n\nSession Chairs: Kento 
 Sato (RIKEN Center for Computational Science (R-CCS)); Chris Schlipalius (
 Pawsey Supercomputing Research Centre; Commonwealth Scientific and Industr
 ial Research Organisation (CSIRO), Australia); and Anja Gerbes (Georg-Augu
 st-Universität Göttingen)\n\n
END:VEVENT
END:VCALENDAR
