BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201803Z
LOCATION:231
DTSTART;TZID=America/Chicago:20251116T164000
DTEND;TZID=America/Chicago:20251116T170000
UID:submissions.supercomputing.org_SC25_sess223_ws_scalah101@linklings.com
SUMMARY:High-Performance and Power-Efficient Emulation of Matrix Multiplic
 ation using INT8 Matrix Engines
DESCRIPTION:Yuki Uchino (RIKEN Center for Computational Science (R-CCS)), 
 Katsuhisa Ozaki (Shibaura Institute of Technology), and Toshiyuki Imamura 
 (RIKEN Center for Computational Science (R-CCS))\n\nRecent architectures i
 ntegrate high-performance and power-efficient matrix engines. \nThese engi
 nes demonstrate remarkable performance in low-precision matrix multiplicat
 ion, which is crucial in deep learning. \nSeveral techniques have been pro
 posed to emulate single- and double-precision general matrix-matrix multip
 lication (SGEMM and DGEMM, respectively) by leveraging such low-precision 
 matrix engines.\nIn this study, we present emulation methods that signific
 antly outperforms conventional approaches.\nOn a GH200 Grace Hopper Superc
 hip, the proposed DGEMM emulation achieves a 1.4x speedup and a 43% improv
 ement in power efficiency compared to native DGEMM for sufficiently large 
 problems.\nThe proposed SGEMM emulation achieves a 3.0x speedup and a 154%
  improvement in power efficiency compared to native SGEMM for sufficiently
  large problems.\nFurthermore, compared to conventional emulation methods,
  the proposed emulation achieves more than 2x higher performance and super
 ior power efficiency.\n\nRecording: Livestreamed, Recorded\n\nRegistration
  Category: Technical Program Reg Pass, Workshop Reg Pass\n\nSession Chairs
 : Vassil Alexandrov (Hartree Centre, STFC); Jack Dongarra (University of T
 ennessee, Knoxville; Oak Ridge National Laboratory (ORNL)); Erik Draeger (
 Lawrence Livermore National Laboratory (LLNL), Center for Applied Scientif
 ic Computing); Philippa Rubin (STFC Hartree Centre); Dieter A. Kranzlmuell
 er (Ludwig-Maxmilians-Universität München, Leibniz Supercomputing Centre (
 LRZ)); and Christian Engelmann (Oak Ridge National Laboratory (ORNL))\n\n
END:VEVENT
END:VCALENDAR
