About the company
At Braze, we have found our people. We’re a genuinely approachable, exceptionally kind, and intensely passionate crew. We seek to ignite that passion by setting high standards, championing teamwork, and creating work-life harmony as we collectively navigate rapid growth on a global scale while striving for greater equity and opportunity – inside and outside our organization.
Responsibilities
- Lead NGINX & Kubernetes Ingress Infrastructure: Architect and operate ingress fleets, configure, tune, and operate high-performance NGINX routing, proxying, and ingress controller layers.
- Own and expand scaling routines for high-throughput, API-based services, leveraging RED metrics, Horizontal Pod Autoscalers (HPA), and customized scaling policies.
- Partner with product teams on resilient architectures: translate product requirements into reliable, highly scalable technology stacks.
- Manage SLIs, SLOs, and error budgets for API services.
- Conduct systems design, bottleneck profiling, and capacity planning.
- Participate in PagerDuty on-call rotation, update runbooks, and lead root-cause analysis and blameless retrospectives.
Requirements
- 5+ years of experience as a DevOps or Site Reliability Engineer in a high-scale production environment.
- Deep hands-on experience with NGINX proxying, routing, and ingress controller layers.
- In-depth proficiency with Kubernetes administration, cluster networking, container orchestration, cluster scheduling, and container deployment.
- Excellent OS-level understanding of Linux/Unix internals.
- Strong programming/scripting skills—Ruby and/or Go preferred (or equivalent languages like Python, Java).
- Experience with IaC technologies such as Terraform, Ansible, Chef, or similar frameworks.
- Strong conceptual understanding of systems design.
- Excellent documentation practices and comfort collaborating asynchronously across remote-first, global engineering teams.
Conditions
- Preferred: familiarity with Redis, Kafka, Postgres, MongoDB; experience with Prometheus, Grafana, Datadog; practical experience with AWS, GCP, Azure.