← Все вакансии/Senior/Cerebras
SeniorOfficeToronto, CAN

Distributed Software Engineer

C
Cerebras
Уровень
Senior
Формат
Office
О роли

Описание вакансии

About the company

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs, delivering industry-leading training and inference speeds. Cerebras works with leading model labs, global enterprises, and AI-native startups.

Responsibilities
  • Declarative, CRD-driven automation of bare-metal networking, OS, and application software across clusters
  • Push-button cluster install, upgrade, and security patching with real downtime budgets, gated by canaries
  • Kubernetes operators that schedule large inference workloads: resource locks, priority queues, network topology, and health-aware placement
  • gRPC control-plane services, authorization, admission webhooks, and quota policy for a multi-tenant fleet
  • Metrics and log pipelines with purpose-built exporters for wafer-scale systems, servers (Redfish, IPMI), and network fabric (gNMI, sFlow), on Prometheus and Grafana
  • Failure detection, HA control planes, and automated recovery, plus CLIs, APIs, and MCP gateway
Requirements
  • 5+ years building and operating production distributed systems or infrastructure software
  • Production-quality Go and Python
  • Real Kubernetes depth: controllers and operators, CRDs, reconciliation semantics, informer caches, admission webhooks, and RBAC
  • Strong debugging skills across distributed systems, Linux, and networking
  • Prometheus and Grafana as a practitioner: PromQL, exporter design, cardinality discipline, useful alerts
  • Strong self-driving capability
  • Demonstrated adoption of AI in engineering workflow
  • Nice to have: bare-metal or HPC fleet operations, scheduler internals, RDMA/RoCE and eBPF networking, Ceph or NVMe-oF, etcd and HA upgrades, inference serving stacks
Conditions
  • Build a breakthrough AI platform beyond the constraints of the GPU
  • Publish and open source cutting-edge AI research
  • Work on one of the fastest AI supercomputers in the world
  • Job stability with startup vitality
  • Non-corporate work culture
Стек и навыки

С чем работаем