← Все вакансии/Senior/jobgether
SeniorRemoteUK

Senior Site Reliability Engineer

J
jobgether
Зарплата
$4,400–$10,000
Уровень
Senior
Формат
Remote
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United Kingdom.

Responsibilities
  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust and maintainable technical solutions.
  • Contribute to platform architecture, infrastructure tooling, technical roadmaps, and engineering priorities, advocating for initiatives that improve reliability and developer experience.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Use operational and incident metrics to identify systemic reliability issues and influence the team's technical strategy.
  • Resolve cross-team infrastructure and platform requests while turning recurring problems into reusable solutions, automation, documentation, and runbooks.
  • Operate and scale production Kubernetes environments and associated container infrastructure.
  • Build and manage cloud infrastructure using AWS or comparable cloud platforms, with strong emphasis on reliability, scalability, and operational efficiency.
  • Develop and maintain infrastructure as code using Terraform and support automated deployment workflows through modern CI/CD platforms.
  • Use AI natively in infrastructure, operations, and development workflows, creating reusable prompts, skills, tooling, and agentic workflows that improve team-wide productivity and reliability.
  • Design infrastructure and systems with AI-assisted engineering in mind, including clean interfaces, strong observability, secure-by-default patterns, CI protections, and review guardrails.
  • Mentor less-senior engineers through actionable feedback, technical guidance, hiring, onboarding, and RFC discussions.
  • Collaborate with Security on infrastructure hardening, threat mitigation, and defensive security practices.
  • Contribute to infrastructure capacity planning, performance optimization, and cost efficiency.
  • Participate in incident response and on-call rotations, helping maintain high standards of platform availability and reliability.
Requirements
  • Solid professional experience in Site Reliability Engineering, DevOps, Platform Engineering, or a closely related discipline.
  • Strong hands-on experience operating and scaling Kubernetes in production, including Docker and the wider container ecosystem.
  • Proven experience designing, building, and managing production cloud infrastructure using AWS or a comparable cloud provider.
  • Strong practical expertise with Terraform and infrastructure-as-code principles.
  • Hands-on experience with reliability engineering frameworks, including SLOs, SLIs, error budgets, alerting, and incident management.
  • Strong observability experience with technologies such as OpenTelemetry, Grafana, Prometheus, or equivalent platforms.
  • Experience designing and operating CI/CD pipelines using GitLab CI, GitHub Actions, or similar technologies.
  • Comfortable working with Golang, Bash, and scripting, with broader programming experience considered an advantage.
  • Demonstrated practical use of AI and agentic workflows in infrastructure, operations, or software engineering, with measurable outcomes beyond simple familiarity with AI tools.
  • Strong understanding of production operations, troubleshooting, automation, and platform reliability.
  • Clear, thoughtful communication skills, particularly in an asynchronous and globally distributed environment.
  • Proactive and curious mindset, with the ability to independently identify problems, take ownership, and drive solutions to completion.
  • Collaborative and respectful approach when working across cultures, time zones, and diverse teams.
  • Experience with a backend programming language such as Elixir, Node.js, Python, or similar is a plus.
  • Experience operating and configuring Linux systems outside cloud environments is beneficial.
  • Security knowledge, including defensive and offensive security concepts, is an advantage.
Conditions
  • 100% remote work, with the ability to work from anywhere.
  • Flexible working hours within an async-first working environment.
  • Flexible paid time off to support a healthy balance between work and personal life.
  • 16 weeks of paid parental leave.
  • Budget for co-working spaces, learning, and wellness, including gym memberships.
  • Mental health support services.
  • Stock options.
  • Home office budget and IT equipment.
  • Competitive, location-aware compensation designed to support fair and equitable pay across global markets.
  • Annual salary range of $53,300–$119,850 USD, with actual compensation determined by location, experience, skills, training, business needs, and market conditions.
  • Significant autonomy and ownership in a globally distributed engineering organization.
  • Opportunities to influence platform architecture, technical strategy, engineering standards, and reliability practices.
Стек и навыки

С чем работаем