BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201220Z
LOCATION:Hall 6
DTSTART;TZID=America/Chicago:20251117T160800
DTEND;TZID=America/Chicago:20251117T160800
UID:submissions.supercomputing.org_SC25_sess553_job138@linklings.com
SUMMARY:HPC Performance and Validation Engineer
DESCRIPTION:The Position\nAs an HPC Validation and Performance Engineer at
  NMC², you will take ownership of the validation and optimization of our H
 PC CPU and GPU calc farms. This critical role will involve developing a va
 lidation and performance baselining framework, which ensures system readin
 ess for AI/ML and HPC workloads across multiple architectures. Your role w
 ill be essential in providing continuous performance benchmarking, real-ti
 me observability, and long-term strategic readiness. You will drive the im
 plementation of advanced tooling and frameworks, maintaining an infrastruc
 ture that is crucial to our cutting-edge research efforts. You will be acc
 ountable for providing data driven performance metrics to support architec
 tural design choices as we continue to globally scale our datacenter footp
 rint. We are looking for someone with deep technical expertise in compute,
  storage or networking optimizations and performance engineering who can d
 evelop solutions that scale with our growing infrastructure. This role dem
 ands a forward-thinking engineer who can anticipate industry trends and ad
 opt emerging architectures and strategies to keep NMC² at the forefront of
  innovation.\n\nResponsibilities:\nArchitecting and implementing a validat
 ion framework to certify the readiness and utilization of GPU nodes across
  a large, distributed HPC environment.\nDefining methodologies to continua
 lly assess performance and optimising infrastructure across AI/ML workload
 s \nDeveloping and executing comprehensive performance testing using indus
 try and customer specific benchmarks, ensuring optimal performance across 
 HPC compute, storage and networking \nContribute to research reports that 
 will describe the discoveries of the benchmarking, evaluating the complete
  HW performance and efficiency \nLeading efforts to debug, identify and th
 en resolve bottlenecks in system performance \nBuilding robust, scalable t
 ools for automated validation and testing, utilizing Python, Go, Kubernete
 s and CI/CD pipelines to streamline continuous validation and benchmarking
  processes \nImplementing monitoring solutions using Prometheus, Grafana a
 nd other modern monitoring technologies to track performance metrics and r
 eal-time health of the cluster \nDefining and implementing best practice f
 or continuous performance validation, ensuring that the infrastructure rem
 ains reliable and efficient as new technologies emerge \nStaying informed 
 on industry trends and advancements to ensure long-term strategic alignmen
 t \nWorking cross-functionally with engineering, infrastructure and resear
 ch teams to align validation efforts with the broader business objectives,
  ensuring that the platform meets evolving research demands\n\nRegistratio
 n Category: Technical Program Reg Pass, Workshop Reg Pass, Tutorial Reg Pa
 ss, Exhibits Reg Pass\n\nCountry: United States of America\n\nCompany: Nor
 thMark Compute and Cloud\n\nIn-Person / Remote: In-person\n\nPart Time / F
 ull Time: Full Time\n\nPosition Type: Permanent\n\n
END:VEVENT
END:VCALENDAR
