About the company
Klaviyo is a company that empowers creators to own their own destiny. The platform is primarily built with Python and React and runs on AWS.
Responsibilities
- Build and operate foundational, security-critical services with a strong emphasis on availability, scalability, latency, and fault tolerance
- Apply software engineering principles to automate infrastructure, reduce operational toil, and improve system reliability at scale
- Design, implement, and evolve systems using SRE best practices
- Define and refine SLIs, SLOs, and error budgets to guide engineering decisions
- Improve observability, alerting, and incident response to reduce mean time to detection and recovery
- Participate in on-call rotations with a focus on sustainable operations and automatic remediations
- Perform quantitative analysis to understand system behavior, capacity constraints, and scaling limits
- Identify systemic risks and reliability bottlenecks and drive long-term, preventative solutions
- Collaborate closely with product, platform, and security engineers to influence architecture early and ship reliable systems
- Mentor and pair with other engineers, helping raise the bar for reliability, operational maturity, and engineering excellence
Requirements
- You write and maintain production-quality code (e.g. Python, Go, or similar) to build internal platforms, automate operations, and improve system reliability
- You have built, deployed, and operated distributed, cloud-native systems and understand failure modes such as partial outages, dependency failures, resource saturation, and cascading impact
- You have experience operating containerized workloads and platforms (e.g. Kubernetes) in production, including deployment strategies, scaling behavior, and service networking
- You are comfortable participating in on-call rotations and diagnosing production issues
- You have designed and operated observability systems and know how to build actionable alerts that reflect real user and service impact
- You apply SRE concepts such as SLIs, SLOs, error budgets, and burn-rate–based alerting to guide engineering decisions and operational response
- You have hands-on experience with infrastructure as code and declarative configuration (e.g. Terraform, Kubernetes manifests, policy-as-code)
- You have performed capacity planning, load testing, and performance analysis for distributed services and platforms
- You routinely contribute to post-incident reviews and drive concrete, code-focused follow-up actions that prevent recurrence
- You are comfortable reviewing and contributing to technical designs, platform APIs, operational runbooks, and system documentation
- You’ve already experimented with AI in work or personal projects
Nice to Have
- Experience supporting security-critical platforms or building internal security tooling
- Familiarity with identity, access management, secrets management, or policy enforcement systems
- Experience operating systems at scale in cloud environments (AWS preferred)
- Background in resilience testing, fault injection, or chaos engineering
- A strong comprehension of algorithms and data structures at scale
Tech Stack
- Python / Django / FastAPI
- MySQL / Redis / Memcached
- RabbitMQ / Celery / Apache Kafka / Apache Pulsar
- AWS / Terraform / Kubernetes