BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260202T201221Z
LOCATION:Hall 6
DTSTART;TZID=America/Chicago:20251117T161800
DTEND;TZID=America/Chicago:20251117T161900
UID:submissions.supercomputing.org_SC25_sess553_job167@linklings.com
SUMMARY:Senior AI-HPC Cluster Engineer - MLOps
DESCRIPTION:NVIDIA has been transforming computer graphics, PC gaming, and
  accelerated computing for more than 25 years. It’s a unique legacy of inn
 ovation that’s fueled by great technology—and amazing people. Today, we’re
  tapping into the unlimited potential of AI to define the next era of comp
 uting. An era in which our GPU acts as the brains of computers, robots, an
 d self-driving cars that can understand the world. Doing what’s never been
  done before takes vision, innovation, and the world’s best talent. As an 
 NVIDIAN, you’ll be immersed in a diverse, supportive environment where eve
 ryone is inspired to do their best work. Come join the team and see how yo
 u can make a lasting impact on the world. NVIDIA reinvented itself over tw
 o decades, inventing the GPU in 1999 and revolutionizing computer graphics
 . Design and implement GPU compute clusters for deep learning and high-per
 formance computing.\n\nWhat you'll be doing:\n\nProvide leadership and str
 ategic mentorship on the management of large-scale HPC systems including t
 he deployment of compute, networking, and storage.\n\nDevelop and improve 
 our ecosystem around GPU-accelerated computing including developing scalab
 le automation solutions.\n\nBuild and nurture customer and cross-team rela
 tionships to consistently support the clusters and address changing user n
 eeds.\n\nSupport our researchers to run their workloads including performa
 nce analysis and optimizations.\n\nConduct root cause analysis and suggest
  corrective action. Proactively find and fix issues before they occur.\n\n
 Build innovative tooling to accelerate researchers' velocity, troubleshoot
 ing, and software performance at scale.\n\nWhat we need to see:\n\nBachelo
 r’s degree in Computer Science, Electrical Engineering or related field or
  equivalent experience.\n\nMinimum of 6 years of experience crafting and o
 perating large scale compute infrastructure.\n\nExperience with AI/HPC job
  schedulers and orchestrators, such as Slurm, K8s or LSF. Applied experien
 ce with AI/HPC workflows that use MPI and NCCL.\n\nProficient in using Lin
 ux including Centos/RHEL and/or Ubuntu Linux distributions. A solid unders
 tanding of container technologies like Enroot, Docker and Podman.\n\nProfi
 ciency in one scripting language (Python, Bash) and at least one compiled 
 language (Golang, Rust, C, C++...).\n\nExperience analyzing and tuning per
 formance for a variety of AI/HPC workloads. Excellent problem-solving to a
 nalyze complex systems, identify bottlenecks, and implement scalable solut
 ions.\n\nExcellent communication and teamwork skills, with the ability to 
 work effectively with diverse teams and individuals.\n\nPassion for contin
 ual learning and staying ahead of new technologies and effective approache
 s in the HPC and AI/ML infrastructure fields.\n\nWays to stand out from th
 e crowd:\n\nExperience with NVIDIA GPUs, CUDA Programming, NCCL and MLPerf
  benchmarking.\n\nExperience with Machine Learning and Deep Learning conce
 pts, algorithms and models.\n\nFamiliarity with High-Speed Networking pert
 aining to HPC including InfiniBand, RDMA, RoCE and Amazon EFA.\n\nUndersta
 nding of fast, distributed storage systems like Lustre and GPFS for AI/HPC
  workload. Experience working with deep learning frameworks including PyTo
 rch, MegatronLM and TensorFlow.\n\nFamiliarity with metrics collection and
  visualization at scale with Prometheus, OpenSearch and Grafana.\n\nNVIDIA
  offers competitive salaries and benefits. Our experienced and talented em
 ployees contribute to our outstanding engineering team's rapid growth. If 
 you're a tech enthusiast, apply now!\n\nYour base salary will be determine
 d based on your location, experience, and the pay of employees in similar 
 positions. The base salary range is 184,000 USD - 287,500 USD for Level 4,
  and 224,000 USD - 356,500 USD for Level 5.\nYou will also be eligible for
  equity and benefits.\n\nRegistration Category: Technical Program Reg Pass
 , Workshop Reg Pass, Tutorial Reg Pass, Exhibits Reg Pass\n\nCountry: Unit
 ed States of America\n\nCompany: NVIDIA Corporation\n\nIn-Person / Remote:
  In-person, Remote\n\nPart Time / Full Time: Full Time\n\nPosition Type: P
 ermanent\n\n
END:VEVENT
END:VCALENDAR
