← Все вакансии/PostHog
RemoteSan Francisco, CA

Site Reliability Engineer

P
PostHog
Формат
Remote
О роли

Описание вакансии

About the company

PostHog is a product analytics platform that started with open-source product analytics and has since shipped more than a dozen products, including PostHog Desktop, a built-in data warehouse, and PostHog AI. We are product-led, default alive, and well-funded. We value transparency, autonomy, shipping fast, time for building, ambition, and being weird.

Responsibilities
  • Operating EKS clusters across several environments with Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments
  • Managing and evolving a multi AWS account organization, provisioning, networking, access control, and cross-account connectivity
  • Maintaining the Terraform/Terragrunt IaC platform - modules, automated plan-on-PR / apply-on-merge pipelines, and safe patterns for shared infrastructure
  • Improving operational tooling around deploys, schema changes, backups, restores, and incident response
  • Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation
  • Optimizing cloud spend as you go
  • Participating in on-call and incident response, with a strong focus on making incidents rarer over time
Requirements
  • Deep ownership of production systems
  • Not afraid of working with stateful infrastructure
  • Love working in AWS, VMs, automation, and making messy systems reliable
  • Enthusiastic drivers: proactive, can fully own projects and get them done
  • Optimistic problem solvers: collaborate, iterate, and ship their way out of anything
  • Grown ups: kind, considerate, professional, low-ego, flexible, respectful
  • Genuine builders: love building stuff
  • Located in US - Pacific timezone
Стек и навыки

С чем работаем