About the company
Zscaler accelerates digital transformation to ensure our customers can be more agile, efficient, resilient, and secure. As an AI-forward enterprise, we are constantly pushing the envelope, leveraging the world’s largest security data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects our customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location.
Responsibilities
- Design and implement highly available, scalable infrastructure across AWS, Azure, GCP, and bare-metal environments.
- Drive an "automation-first" culture by writing code (Python/Go) to eliminate manual toil and build self-healing systems.
- Implement and maintain sophisticated observability (Prometheus, Grafana, OpenTelemetry), define SLIs/SLOs, and establish error budgets.
- Act as a lead Incident Commander (TDO on-call), develop response playbooks, and conduct deep-dive post-incident analyses.
- Partner with Engineering and partner teams to conduct operability reviews.
Requirements
- Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions.
- 4 years of experience managing reliability, scalability, and availability for large-scale production services with deep expertise in programming (e.g., Python, Go, or C/C++).
- Strong background in networking protocols, Linux/FreeBSD systems, and distributed architecture.
- Experience in high-stakes incident management and participation in a 24/7 on-call rotation.
- Proficiency in leveraging ITIL frameworks and incident data to drive service maturity.
Conditions
- Base pay range: $103,600 – $148,000 USD.