← Все вакансии/Senior/Braze
SeniorOfficeVancouver

Site Reliability Engineer

B
Braze
Уровень
Senior
Формат
Office
О роли

Описание вакансии

About the company

At Braze, we have found our people. We’re a genuinely approachable, exceptionally kind, and intensely passionate crew. We seek to ignite that passion by setting high standards, championing teamwork, and creating work-life harmony as we collectively navigate rapid growth on a global scale while striving for greater equity and opportunity – inside and outside our organization.

Responsibilities
  • Lead NGINX & Kubernetes Ingress Infrastructure: Architect and operate ingress fleets, configure, tune, and operate high-performance NGINX routing, proxying, and ingress controller layers.
  • Own and expand scaling routines for high-throughput, API-based services, leveraging RED metrics, Horizontal Pod Autoscalers (HPA), and customized scaling policies.
  • Partner with product teams on resilient architectures: translate product requirements into reliable, highly scalable technology stacks.
  • Manage SLIs, SLOs, and error budgets for API services.
  • Conduct systems design, bottleneck profiling, and capacity planning.
  • Participate in PagerDuty on-call rotation, update runbooks, and lead root-cause analysis and blameless retrospectives.
Requirements
  • 5+ years of experience as a DevOps or Site Reliability Engineer in a high-scale production environment.
  • Deep hands-on experience with NGINX proxying, routing, and ingress controller layers.
  • In-depth proficiency with Kubernetes administration, cluster networking, container orchestration, cluster scheduling, and container deployment.
  • Excellent OS-level understanding of Linux/Unix internals.
  • Strong programming/scripting skills—Ruby and/or Go preferred (or equivalent languages like Python, Java).
  • Experience with IaC technologies such as Terraform, Ansible, Chef, or similar frameworks.
  • Strong conceptual understanding of systems design.
  • Excellent documentation practices and comfort collaborating asynchronously across remote-first, global engineering teams.
Conditions
  • Preferred: familiarity with Redis, Kafka, Postgres, MongoDB; experience with Prometheus, Grafana, Datadog; practical experience with AWS, GCP, Azure.
Стек и навыки

С чем работаем