← Все вакансии/Lead/Cloudlinux
LeadRemoteYerevan

Lead Site Reliability Engineer

C
Cloudlinux
Уровень
Lead
Формат
Remote
О роли

Описание вакансии

About the company

Cloudlinux is a company that provides security solutions for Linux servers, including Imunify360, a multi-layer security suite.

Responsibilities
  • Define what "working" means for ~70 components: run SLI definition with squad leads and senior engineers, build taxonomy (Service SLIs, Fleet SLIs, Control-efficacy SLIs, Delivery SLIs, Pipeline SLIs), attach SLO, error budget, and owning squad to each.
  • Build the collection system: design and build pipeline to get indicators off the fleet into a queryable store, extend agent-side and service-side instrumentation in Python, Go, and Rust, consolidate dashboards and reporting paths.
  • Build alerting and alert management: symptom-based, SLO-anchored alerting with multi-window burn-rate semantics, three-tier taxonomy (page/ticket/dashboard), ensure every alert has owner, runbook, and documented failure mode.
  • Build escalation: component-to-owning-squad ownership map, severity matrix, acknowledgement SLAs, follow-the-sun rota design, incident command practice and blameless postmortems.
Requirements
  • Substantial production-engineering or SRE experience, including defining SLO framework.
  • Strong Python, comfortable reading and modifying Go or Rust.
  • Deep practical grip on time-series and event telemetry at scale: Prometheus/OpenMetrics, Grafana, Alertmanager-class routing, columnar store (ClickHouse or equivalent).
  • Distributed systems debugging on bare metal and long-lived hosts.
  • Configuration management and CI at production scale: Ansible, GitLab CI, Jenkins or close equivalents.
  • Judgement to design measurement for machines you do not own and cannot scrape.
Стек и навыки

С чем работаем