Close

Presentation

Staff Production Engineer
·
Groq
·
Remote, USA
DescriptionMission:
Join the team that builds and operates Groq’s real-time, distributed inference system delivering large-scale inference for LLMs and next-gen AI applications at ultra-low latency. As a Low-Level Production Engineer, your mission is to ensure reliability, fault tolerance, and operational excellence in Groq’s LPU-powered infrastructure. You’ll work deep in the stack—bridging distributed runtime systems with the hardware—to keep Groq systems fast, stable, and production-ready at scale.

Responsibilities & opportunities in this role:
Production Reliability: Operate and harden Groq’s distributed runtime across thousands of LPUs, ensuring uptime and resilience under dynamic global workloads.
Low-Level Debugging: Diagnose and resolve hardware-software integration issues in live environments, from datacenter level events to single component failures.
Observability & Diagnostics: Build tools and infrastructure to improve real-time system monitoring, fault detection, and SLO tracking.
Automation & Scale: Automate deployment workflows, failover systems, and operational playbooks to reduce overhead and accelerate reliability improvements.
Performance & Optimization: Profile and tune production systems for throughput, latency, and determinism—every cycle counts.
Cross-Functional Collaboration: Partner with compiler, hardware, infra, and data center teams to deliver robust, fault-tolerant production systems.

Ideal candidates have/are:
Proven experience in production engineering across the stack and operating large-scale distributed systems.
Deep knowledge of computer architecture, operating systems, and hardware-software interfaces.
Skilled in low-level systems programming (C/C++ or Rust), with scripting fluency (Python, Bash, or Go).
Comfortable debugging complex issues close to the metal—kernels, firmware, or hardware-aware code paths.
Strong background in automation, CI/CD, and building reliable systems that scale.
Thrive across environments—from kernel internals to distributed runtimes to data center operations.
Communicate clearly, make pragmatic decisions, and take ownership of long-term outcomes.

Nice to have:
Experience operating high-performance, real-time systems at scale (ML inference, HPC, or similar).
Familiarity with GPUs, FPGAs, or ASICs in production environments.
Prior exposure to ML frameworks (e.g., PyTorch) or compiler tooling (e.g., MLIR).
Track record of delivering complex production systems in high-impact environments.

Attributes of a Groqster:
Humility – Egos are checked at the door
Collaborative & Team Savvy – We make up the smartest person in the room, together
Growth & Giver Mindset – Learn it all versus know it all, we share knowledge generously
Curious & Innovative – Take a creative approach to projects, problems, and design
Passion, Grit, & Boldness – No-limit thinking, fueling informed risk taking
Company DescriptionGroq delivers fast, efficient AI inference. Our LPU-based system powers GroqCloud™, giving businesses and developers the speed and scale they need. From our Bay Area roots to our growing global presence, we are on a mission to make high performance AI compute more accessible and affordable. When real-time AI is within reach, anything is possible. Build fast.
·
·
Event Type
Job Posting
TimeMonday, 17 November 20254:51pm - 4:51pm CST
LocationHall 6
Countries
United States of America
Companies
Groq
In-Person / Remotes
Remote
Part Time / Full Times
Full Time
Position Types
Permanent