About the company
Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster.
Responsibilities
- Hire, manage, mentor, and coach a geographically distributed team of infrastructure and SRE engineers across SF Bay Area, India, and Europe, helping them grow both technically and professionally.
- Drive and uphold a culture of respect, integrity, inclusion, ownership, and accountability within the team.
- Set clear goals and provide regular feedback to ensure your team is motivated and aligned with the platform and company vision.
- Build a high-performing team with the skills and practices to reliably operate and evolve cloud agnostic infrastructure at scale.
- Foster psychological safety and cross-regional collaboration across time zones.
- Own the architecture and evolution of Postman's cloud agnostic infrastructure, driving improvements that increase reliability, performance, and cost efficiency at scale.
- Lead the design and implementation of infrastructure improvements across Kubernetes, Cluster API, Argo, Helm, Crossplane, service mesh (Istio), AWS, and Azure environments.
- Own the SRE function end to end: SLIs/SLOs, error budgets, capacity planning, incident management, and reliability engineering practices across the platform.
- Set the technical direction for how Postman's infrastructure and reliability practices evolve to support a large engineering organization with hundreds of services and dozens of teams, with a focus on enabling product teams to operate autonomously.
- Partner with platform, security, and product engineering teams to ensure infrastructure and reliability decisions align with broader company goals.
- Ensure infrastructure and reliability best practices are upheld across the organization, including GitOps, CI/CD, observability, on-call, and incident response.
- Own the infrastructure and reliability roadmap, balancing operational reliability with longer-term architectural investments.
- Break down complex infrastructure and reliability initiatives into clear, actionable milestones and manage delivery on time and at high quality.
- Proactively identify and resolve roadblocks, working across teams to unblock engineering work and minimize customer impact.
- Work closely with stakeholders across engineering, product, and security teams to align on infrastructure and reliability priorities and constraints.
- Foster open communication within and across teams, promoting transparency on system health, risk, error budgets, and roadmap.
- Represent infrastructure and SRE in leadership forums, advocating for technical needs and communicating clearly on trade-offs.
- Own Postman's SRE function, establishing and evolving on-call practices, escalation policies, monitoring, and incident management across the company.
- Define and track SLIs/SLOs and error budgets for critical services, using them to guide investment decisions and prioritization.
- Drive a blameless postmortem culture, ensuring incidents produce durable learnings, clear action items, and measurable follow-through.
- Drive a culture of continuous improvement, learning from incidents, automating toil, and reducing operational burden so that product engineering teams can ship independently.
- Champion proactive reliability practices such as load testing and capacity planning to stay ahead of scale.
- Maintain high standards for security, cost management, and infrastructure quality across all environments.
Requirements
- You are a seasoned infrastructure and reliability leader with deep technical expertise.