About the company
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs, delivering industry-leading training and inference speeds. Cerebras works with leading model labs, global enterprises, and AI-native startups.
Responsibilities
- Declarative, CRD-driven automation of bare-metal networking, OS, and application software across clusters
- Push-button cluster install, upgrade, and security patching with real downtime budgets, gated by canaries
- Kubernetes operators that schedule large inference workloads: resource locks, priority queues, network topology, and health-aware placement
- gRPC control-plane services, authorization, admission webhooks, and quota policy for a multi-tenant fleet
- Metrics and log pipelines with purpose-built exporters for wafer-scale systems, servers (Redfish, IPMI), and network fabric (gNMI, sFlow), on Prometheus and Grafana
- Failure detection, HA control planes, and automated recovery, plus CLIs, APIs, and MCP gateway
Requirements
- 5+ years building and operating production distributed systems or infrastructure software
- Production-quality Go and Python
- Real Kubernetes depth: controllers and operators, CRDs, reconciliation semantics, informer caches, admission webhooks, and RBAC
- Strong debugging skills across distributed systems, Linux, and networking
- Prometheus and Grafana as a practitioner: PromQL, exporter design, cardinality discipline, useful alerts
- Strong self-driving capability
- Demonstrated adoption of AI in engineering workflow
- Nice to have: bare-metal or HPC fleet operations, scheduler internals, RDMA/RoCE and eBPF networking, Ceph or NVMe-oF, etcd and HA upgrades, inference serving stacks
Conditions
- Build a breakthrough AI platform beyond the constraints of the GPU
- Publish and open source cutting-edge AI research
- Work on one of the fastest AI supercomputers in the world
- Job stability with startup vitality
- Non-corporate work culture