About the company
This position is listed on behalf of a partner company, who manages all applications and next steps.
Responsibilities
- Lead and mentor Site Reliability Engineering initiatives focused on highly available managed gateway solutions.
- Architect, develop, and maintain resilient cloud-native systems using technologies such as Kubernetes, Golang, and major public cloud platforms.
- Own the complete operational lifecycle, including monitoring, alerting, incident response, troubleshooting, and continuous service improvement.
- Establish and improve automation, self-service tooling, and deployment workflows to enhance engineering efficiency.
- Define, monitor, and optimize service level objectives (SLOs) and service level indicators (SLIs) to maintain platform reliability.
- Drive best practices around architecture, technical debt management, scalability, and operational excellence.
- Collaborate with product, engineering, and customer-facing teams to improve platform capabilities and operational readiness.
- Partner directly with enterprise customers to support onboarding, implementation, and successful production deployments.
- Analyze complex customer environments across AWS, GCP, and Azure, adapting solutions to unique infrastructure requirements.
- Act as a senior technical escalation point for complex customer implementations and reliability challenges.
- Transform recurring implementation patterns into scalable solutions, documentation, and reusable platform capabilities.
- Provide product feedback based on customer requirements, operational insights, and real-world deployment experiences.
Requirements
- Extensive experience working as a Site Reliability Engineer, DevOps Engineer, or similar role focused on highly available and distributed systems.
- Deep knowledge of Kubernetes, cloud-native architectures, and modern infrastructure practices.
- Experience working across multiple cloud providers such as AWS, GCP, and Azure.
- Strong programming skills in Golang or similar modern languages used for automation and tooling.
- Proven experience building and maintaining CI/CD pipelines and infrastructure-as-code solutions using tools such as Terraform or Ansible.
- Strong understanding of monitoring, logging, and observability platforms including Prometheus, Grafana, ELK, Datadog, or similar technologies.
- Experience with managed services, API gateways, networking infrastructure, or related technologies is highly valuable.
- Strong ownership mindset with the ability to manage critical systems and drive effective incident resolution.
- Excellent collaboration and communication skills with the ability to work across distributed teams and time zones.
- Ability to operate effectively in a fast-paced environment where priorities evolve quickly.
- Preferred Qualifications:
- Experience with service mesh technologies such as Istio or Linkerd.
- Familiarity with database administration for high-throughput systems, including PostgreSQL or Cassandra.
- Contributions to open-source infrastructure or reliability engineering projects.
- Relevant cloud certifications such as AWS Certified DevOps Engineer or Certified Kubernetes Administrator (CKA).
Conditions
- Competitive compensation range of CA$118K–CA$167K depending on location, experience, and qualifications.
- Fully remote work opportunity within Canada.
- Opportunity to work on large-scale cloud infrastructure and enterprise technology solutions.
- Exposure to complex multi-cloud environments and cutting-edge reliability engineering practices.
- Collaborative environment with globally distributed engineering and product teams.
- Opportunity to influence platform strategy and technical direction.
- Professional growth opportunities through challenging projects and continuous learning.
- Supportive culture focused on innovation, ownership, and engineering excellence.