About the company
Cloudlinux is a company that provides security solutions for Linux servers, including Imunify360, a multi-layer security suite.
Responsibilities
- Define what "working" means for ~70 components: run SLI definition with squad leads and senior engineers, build taxonomy (Service SLIs, Fleet SLIs, Control-efficacy SLIs, Delivery SLIs, Pipeline SLIs), attach SLO, error budget, and owning squad to each.
- Build the collection system: design and build pipeline to get indicators off the fleet into a queryable store, extend agent-side and service-side instrumentation in Python, Go, and Rust, consolidate dashboards and reporting paths.
- Build alerting and alert management: symptom-based, SLO-anchored alerting with multi-window burn-rate semantics, three-tier taxonomy (page/ticket/dashboard), ensure every alert has owner, runbook, and documented failure mode.
- Build escalation: component-to-owning-squad ownership map, severity matrix, acknowledgement SLAs, follow-the-sun rota design, incident command practice and blameless postmortems.
Requirements
- Substantial production-engineering or SRE experience, including defining SLO framework.
- Strong Python, comfortable reading and modifying Go or Rust.
- Deep practical grip on time-series and event telemetry at scale: Prometheus/OpenMetrics, Grafana, Alertmanager-class routing, columnar store (ClickHouse or equivalent).
- Distributed systems debugging on bare metal and long-lived hosts.
- Configuration management and CI at production scale: Ansible, GitLab CI, Jenkins or close equivalents.
- Judgement to design measurement for machines you do not own and cannot scrape.