← Все вакансии/Senior/Klaviyo
SeniorHybridDublin

Software Engineer

K
Klaviyo
Уровень
Senior
Формат
Hybrid
О роли

Описание вакансии

About the company

Klaviyo is a company that empowers creators to own their own destiny. The platform is primarily built with Python and React and runs on AWS.

Responsibilities
  • Build and operate foundational, security-critical services with a strong emphasis on availability, scalability, latency, and fault tolerance
  • Apply software engineering principles to automate infrastructure, reduce operational toil, and improve system reliability at scale
  • Design, implement, and evolve systems using SRE best practices
  • Define and refine SLIs, SLOs, and error budgets to guide engineering decisions
  • Improve observability, alerting, and incident response to reduce mean time to detection and recovery
  • Participate in on-call rotations with a focus on sustainable operations and automatic remediations
  • Perform quantitative analysis to understand system behavior, capacity constraints, and scaling limits
  • Identify systemic risks and reliability bottlenecks and drive long-term, preventative solutions
  • Collaborate closely with product, platform, and security engineers to influence architecture early and ship reliable systems
  • Mentor and pair with other engineers, helping raise the bar for reliability, operational maturity, and engineering excellence
Requirements
  • You write and maintain production-quality code (e.g. Python, Go, or similar) to build internal platforms, automate operations, and improve system reliability
  • You have built, deployed, and operated distributed, cloud-native systems and understand failure modes such as partial outages, dependency failures, resource saturation, and cascading impact
  • You have experience operating containerized workloads and platforms (e.g. Kubernetes) in production, including deployment strategies, scaling behavior, and service networking
  • You are comfortable participating in on-call rotations and diagnosing production issues
  • You have designed and operated observability systems and know how to build actionable alerts that reflect real user and service impact
  • You apply SRE concepts such as SLIs, SLOs, error budgets, and burn-rate–based alerting to guide engineering decisions and operational response
  • You have hands-on experience with infrastructure as code and declarative configuration (e.g. Terraform, Kubernetes manifests, policy-as-code)
  • You have performed capacity planning, load testing, and performance analysis for distributed services and platforms
  • You routinely contribute to post-incident reviews and drive concrete, code-focused follow-up actions that prevent recurrence
  • You are comfortable reviewing and contributing to technical designs, platform APIs, operational runbooks, and system documentation
  • You’ve already experimented with AI in work or personal projects
Nice to Have
  • Experience supporting security-critical platforms or building internal security tooling
  • Familiarity with identity, access management, secrets management, or policy enforcement systems
  • Experience operating systems at scale in cloud environments (AWS preferred)
  • Background in resilience testing, fault injection, or chaos engineering
  • A strong comprehension of algorithms and data structures at scale
Tech Stack
  • Python / Django / FastAPI
  • MySQL / Redis / Memcached
  • RabbitMQ / Celery / Apache Kafka / Apache Pulsar
  • AWS / Terraform / Kubernetes
Стек и навыки

С чем работаем