About the company
This position is listed on behalf of a partner company, who manages all applications and next steps.
Responsibilities
- Lead and develop a team of engineers responsible for production cloud infrastructure, fostering a strong engineering culture and effective execution.
- Guide the design, implementation, and evolution of Kubernetes-based, multi-cloud infrastructure environments.
- Ensure infrastructure solutions meet high standards for reliability, scalability, security, and operational excellence.
- Collaborate with product engineering, security, support, and engineering leadership teams to enable safe and efficient software delivery.
- Improve infrastructure automation, developer tooling, and platform capabilities to support growth without unnecessary complexity.
- Establish effective approaches for using AI in infrastructure operations, including automation, incident analysis, documentation, and operational improvements.
- Balance immediate operational needs with long-term investments in platform evolution, cloud expansion, and infrastructure innovation.
- Coach engineers, support professional growth, manage priorities, and create an environment where teams can perform at their best.
- Drive improvements in observability, incident response, reliability practices, and infrastructure decision-making processes.
Requirements
- 3–5 years of experience leading engineers who build and operate production infrastructure for SaaS products.
- Strong technical background in cloud infrastructure engineering, platform engineering, Site Reliability Engineering (SRE), or related disciplines.
- Hands-on production experience with Kubernetes, including scaling, reliability management, and cluster lifecycle operations.
- Experience working with multiple major cloud providers and understanding the challenges and trade-offs of multi-cloud environments.
- Proven ability to collaborate with senior engineers and cross-functional stakeholders to deliver major infrastructure improvements.
- Strong people leadership skills, including coaching, performance management, prioritization, and team development.
- Experience improving infrastructure reliability, scalability, security, and automation practices.
- Clear understanding of how AI can support infrastructure engineering workflows, with good judgment around responsible and effective adoption.
- Experience with infrastructure-as-code, policy-as-code, platform automation, observability, incident management, or cloud cost optimization is a plus.
- Experience managing global or multi-region SaaS infrastructure is considered an advantage.
Conditions
- Competitive, location-based compensation package with transparent salary structures.
- Opportunity to receive equity participation through a company stock option program where applicable.
- Fully remote working environment with flexibility to work from anywhere.
- Home office budget and coworking space support to create an effective workspace.
- Flexible time off policies designed to support work-life balance.
- Opportunities to connect with colleagues through annual gatherings, team offsites, and community events.
- Comprehensive health and wellbeing support, including mental health resources and lifestyle benefits where available.
- Retirement and pension plans, along with life and disability insurance depending on location.
- Annual learning and development budget for courses, certifications, books, and professional growth.
- Opportunity to work on impactful cloud infrastructure challenges within a global technology organization.