About the company
At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements. If you’re a close but not exact match with the description, we hope you’ll still consider applying.
Responsibilities
- Contribute to the architecture and evolution of backend services that power product recommendations across Klaviyo experiences (email, SMS, KAgent, onsite, etc.), meeting standards for reliability, performance, and clear APIs.
- Contribute to and maintain robust, large-scale data processing pipelines (e.g., using Apache Spark or similar frameworks) that transform raw events and catalog data into high-quality features and inputs for recommendation models, ensuring data quality and lineage.
- Collaborate closely with ML engineers and product stakeholders to productionize recommendation models—defining high-level interfaces, feature contracts, and deployment patterns for batch and/or real-time inference systems.
- Contribute to the development of the vector database that powers recommendation, semantic search, and agentic use cases.
- Ensure data and service observability (metrics, logging, tracing, dashboards) to facilitate recommendations that are correct, explainable, fast, and highly available for all customers.
- Work with Product to break down projects into clear milestones, balancing the need for rapid experimentation with technical soundness and long-term maintainability.
- Lead data-driven decision making and A/B testing efforts—ensuring recommendation systems are instrumented with the right metrics, and independently interpreting results to guide future product and engineering iterations.
- Participate in on-call and incident response for the systems you own, driving major post-incident follow-ups that substantially improve the resilience and operability of our recommendation stack.
- Integrate AI into your and the team’s development workflow from the ground up—for example, using AI to accelerate development, automate complex tests, or build smarter monitoring and debugging tools.
- Share knowledge, mentor junior engineers, and define best practices on working with large-scale data frameworks, distributed systems, and integrating ML into production systems.
Requirements
- 2+ years of professional software engineering experience with a focus on backend and distributed systems at scale; you have a proven track record working on production services and optimizing for latency, reliability, and operability as well as business requirements.
- Proficient in Python and open to working in other languages.
- Comfortable with cloud-native architectures (AWS preferred) and container orchestration (e.g., Kubernetes); you manage infrastructure and CI/CD pipelines as a core part of your development process.
- Experience in data-driven decision making and A/B testing—you can define (or are interested in learning how to) how to instrument experiments, read and interpret results, and ensure learnings are folded back into system design.
- Comfortable designing and querying data models in relational, analytical, and NoSQL datastores (e.g., Postgres, MySQL, data warehouses, Redis, vector databases).
- Feel at home with modern DevOps practices (CI/CD, monitoring, alerting) and how to apply them to architect large-scale data and recommendation systems.
- Track record of owning features end-to-end—from initial technical design and implementation through rollout, monitoring, and sustained iteration.
- Excellent technical collaborator and communicator: you can clearly articulate complex technical trade-offs to both technical peers and non-technical partners, and you work effectively to drive alignment across ML Engineers, Software Engineers, PMs, and other teams.
- You are a self-starter who has actively experimented with AI in work or personal projects and are excited to responsibly explore and define new AI tools and workflows to enhance team productivity and system intelligence.
Conditions
- Onsite 5x a week in Boston, MA.