About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United Kingdom.
Responsibilities
- Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust and maintainable technical solutions.
- Contribute to platform architecture, infrastructure tooling, technical roadmaps, and engineering priorities, advocating for initiatives that improve reliability and developer experience.
- Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
- Use operational and incident metrics to identify systemic reliability issues and influence the team's technical strategy.
- Resolve cross-team infrastructure and platform requests while turning recurring problems into reusable solutions, automation, documentation, and runbooks.
- Operate and scale production Kubernetes environments and associated container infrastructure.
- Build and manage cloud infrastructure using AWS or comparable cloud platforms, with strong emphasis on reliability, scalability, and operational efficiency.
- Develop and maintain infrastructure as code using Terraform and support automated deployment workflows through modern CI/CD platforms.
- Use AI natively in infrastructure, operations, and development workflows, creating reusable prompts, skills, tooling, and agentic workflows that improve team-wide productivity and reliability.
- Design infrastructure and systems with AI-assisted engineering in mind, including clean interfaces, strong observability, secure-by-default patterns, CI protections, and review guardrails.
- Mentor less-senior engineers through actionable feedback, technical guidance, hiring, onboarding, and RFC discussions.
- Collaborate with Security on infrastructure hardening, threat mitigation, and defensive security practices.
- Contribute to infrastructure capacity planning, performance optimization, and cost efficiency.
- Participate in incident response and on-call rotations, helping maintain high standards of platform availability and reliability.
Requirements
- Solid professional experience in Site Reliability Engineering, DevOps, Platform Engineering, or a closely related discipline.
- Strong hands-on experience operating and scaling Kubernetes in production, including Docker and the wider container ecosystem.
- Proven experience designing, building, and managing production cloud infrastructure using AWS or a comparable cloud provider.
- Strong practical expertise with Terraform and infrastructure-as-code principles.
- Hands-on experience with reliability engineering frameworks, including SLOs, SLIs, error budgets, alerting, and incident management.
- Strong observability experience with technologies such as OpenTelemetry, Grafana, Prometheus, or equivalent platforms.
- Experience designing and operating CI/CD pipelines using GitLab CI, GitHub Actions, or similar technologies.
- Comfortable working with Golang, Bash, and scripting, with broader programming experience considered an advantage.
- Demonstrated practical use of AI and agentic workflows in infrastructure, operations, or software engineering, with measurable outcomes beyond simple familiarity with AI tools.
- Strong understanding of production operations, troubleshooting, automation, and platform reliability.
- Clear, thoughtful communication skills, particularly in an asynchronous and globally distributed environment.
- Proactive and curious mindset, with the ability to independently identify problems, take ownership, and drive solutions to completion.
- Collaborative and respectful approach when working across cultures, time zones, and diverse teams.
- Experience with a backend programming language such as Elixir, Node.js, Python, or similar is a plus.
- Experience operating and configuring Linux systems outside cloud environments is beneficial.
- Security knowledge, including defensive and offensive security concepts, is an advantage.
Conditions
- 100% remote work, with the ability to work from anywhere.
- Flexible working hours within an async-first working environment.
- Flexible paid time off to support a healthy balance between work and personal life.
- 16 weeks of paid parental leave.
- Budget for co-working spaces, learning, and wellness, including gym memberships.
- Mental health support services.
- Stock options.
- Home office budget and IT equipment.
- Competitive, location-aware compensation designed to support fair and equitable pay across global markets.
- Annual salary range of $53,300–$119,850 USD, with actual compensation determined by location, experience, skills, training, business needs, and market conditions.
- Significant autonomy and ownership in a globally distributed engineering organization.
- Opportunities to influence platform architecture, technical strategy, engineering standards, and reliability practices.