Voice AI platform powering 1 billion calls for companies like Amazon Ring, Intuit, ServiceTitan, and New York Life
Trusted by 1 million developers building voice agents
Backed by Peak XV, Bessemer, Kleiner Perkins, M12, Y Combinator, with $72M raised
Responsibilities
Learn Vapi's architecture, production environment, incident history, and reliability practices; contribute an initial operational or reliability improvement
Own a reliability workstream across observability, incident response, capacity, performance, or production automation
Reduce manual work and improve how the team detects, understands, and responds to failures
Become the go-to owner for a meaningful part of Vapi's reliability surface
Deliver durable improvements to failure prevention or recovery and propose a roadmap for next reliability investments
Requirements
Senior or staff-level software engineer with meaningful SRE, production engineering, or infrastructure experience in distributed systems
Write production-quality software and have built reliability tooling or automation yourself
Deep experience with observability, incident response, failure analysis, capacity, and production system health practices
Comfortable with Kubernetes, networking, and cloud infrastructure; able to debug across application and infrastructure boundaries
Clear reasoning about failure modes and balancing reliability investments with product and engineering velocity
Experience with real-time networking or telephony, Envoy, Postgres, Redis, Kafka, Aurora, ClickHouse, or a Google-style SRE environment is a strong plus
Conditions
Base salary of $280,000 to $314,000 and excellent equity ownership
Comprehensive health coverage: medical, dental, and vision plans
Quarterly off-sites
Flexible time off
Catered meals, transportation, gym, and a $10k annual L&D budget