Close

Presentation

Senior Staff Linux Systems Engineer, Compute & Storage
·
Groq
·
Remote, USA
DescriptionMission:
At Groq, we are building a custom cloud from the ground up - one data center at a time. Our Compute Storage team owns the systems that turn racks of bare metal into production-ready Kubernetes clusters powering the next generation of AI workloads.

We are looking for a Sr. Staff Linux Systems Engineer to help us scale this effort. This role focuses on creating a reliable, performant and secure foundation for the Groq Cloud. You will work with your infrastructure peers to enable and optimize compute nodes and storage clusters that form Groq Cloud. We're looking for someone passionate about infrastructure who enjoys debugging close to the metal. If you're eager to grow your skills in deploying, scaling, and optimizing bare metal to support complex distributed HPC in the expanding inference market – we would love to talk.

Responsibilities & opportunities in this role:
Kernel and OS level enablement and optimization for compute nodes (GPU, LPU) and storage clusters.
Work with infrastructure peers to define optimal health standards for all production servers, including certified OS, Kernel, BIOS/FW versions.
Strengthen security posture through improving system level CVE response
Debug and resolve systems level performance and reliability issues in the fleet.
Work with vendors to debug and resolve BIOS/FW issues.
Support design and deployment of large GPU clusters.
Lead cross-functional collaboration with data center operations, networking, and platform teams to ensure infrastructure is fully integrated and production-ready.
Follow best practices and standards for infrastructure-as-code and configuration management using Git, Flux, Terraform, and related tools.
Set technical direction and maintain high-quality system documentation, operational runbooks, and internal tooling that improve the resilience, repeatability, and observability of the infrastructure stack.

Ideal candidates have/are:
Experience with Linux OS management in large virtualized environments.
Deep Kernel knowledge with experience working with the upstream community to resolve bugs.
Experience deploying large GPU clusters with network fabric.
Familiarity with infrastructure-as-code and Git-based workflows (e.g., Terraform, Flux, Kustomize).
Ability to write and maintain basic tooling in Go, Python, or Bash.
Understanding of networking fundamentals (IPAM, VLANs, DHCP, DNS).
Working knowledge of storage concepts (block vs object, NFS, RAID, etc.).
Strong sense of ownership and a willingness to dive into hardware, firmware, or low-level provisioning issues.

Nice to Have:
Exposure to Talos Linux.
Experience with maintaining a production Kubernetes environment.
Hardware SKU definition and lifecycle management.

Attributes of a Groqster:
Humility - Egos are checked at the door
Collaborative & Team Savvy - We make up the smartest person in the room, together
Growth & Giver Mindset - Learn it all versus know it all, we share knowledge generously
Curious & Innovative - Take a creative approach to projects, problems, and design
Passion, Grit, & Boldness - no limit thinking, fueling informed risk taking
Company DescriptionGroq delivers fast, efficient AI inference. Our LPU-based system powers GroqCloud™, giving businesses and developers the speed and scale they need. From our Bay Area roots to our growing global presence, we are on a mission to make high performance AI compute more accessible and affordable. When real-time AI is within reach, anything is possible. Build fast.
·
·
Event Type
Job Posting
TimeMonday, 17 November 20254:50pm - 4:50pm CST
LocationHall 6
Countries
United States of America
Companies
Groq
In-Person / Remotes
Remote
Part Time / Full Times
Full Time
Position Types
Permanent